How Do You Know If an AI Is Coaching?

An AI that coaches should move learners from asking questions to trying things. We built a way to measure whether Coach does this, and we're sharing what we’ve found so far.

Coach, the AI-powered career coaching tool we’ve been building since 2023, is designed to coach rather than just answer questions. For most of the time we've been building it, that has been a conviction we designed around, but not something we could easily measure. We were fairly confident it was doing more than answering well, but we couldn't point to evidence that it was.

Measuring it is harder than it sounds, and most of the difficulty lives in the question itself: what would an AI be doing differently if it were coaching instead of answering, and could we see it in a transcript? Over the past few months we tried to work that out. This post walks through how we made coaching measurable and what those measurements showed, including early signs of success and areas for further improvement.

What coaching looks like, concretely

Start with the situation that Coach's ‘Explore career paths’ activity is built for: a learner who doesn't know where to begin.

A tool optimized to answer treats that as a request for options. Ask it some version of "what career should I consider," and it returns a clean list of fields, maybe with a suggestion to take an interest assessment. The learner ends up with more possibilities than they came in with, and no better sense of which one is theirs.

A tool optimized to coach treats the not-knowing as the thing to work on. Coach doesn't open with careers. It acknowledges that the choice feels overwhelming and offers a set of skills the learner can build instead: recognizing their own strengths and values, weighing options, finding people who can help, testing whether a path fits by taking a first step. The learner chooses where to start. The question shifts from "what should I do" to "how do I work this out," which is something a person can make progress on.

The ‘Explore career paths’ activity is now designed to help learners make career development progress, instead of simply providing a list of occupations to consider.

The difference that matters is between handing someone options and giving them a way to navigate toward one. That difference is what we set out to measure.

The framework guiding our approach

There's no single feature that turns an answer engine into a coach. Getting Coach to behave like one is a portfolio of investments we keep adding to. Some of it is personalization: adapting to where a learner is and what they've already tried, including the surprisingly hard problem of getting Coach to ask before it tells. Evaluation is another piece, a tiered system that watches for quality regressions night to night and runs safety checks before anything ships. And much of it is human, with user interviews, partner pilots, and structured review from career-development experts all feeding back into how the product changes. The part worth going deep on here is the one that makes the rest measurable in the first place: the learning-science framework underneath Coach's activities.

The framework we’re applying is called KAR, for Knowledge, Application, and Reflection: a career-adapted condensation of Bloom's taxonomy developed by the National Career Development Association. Knowledge is learning about something: what networking is, what a resume is for. Application is doing it: drafting the outreach message, revising the resume, shadowing a professional. Reflection is making sense of the experience afterward: what worked, what you'd change next time.

Knowledge Learning about it what networking is
Application Doing it drafting the message
Reflection Making sense of it what you'd change
The KAR progression: Knowledge means learning about something, Application means doing it, then Reflection means making sense of it afterward.

KAR matters to us because of a finding that has held up across decades of research: learners who reach Application and Reflection tend to see better outcomes than learners who stay in Knowledge. The research on how people actually use chatbots points the other way. Most users stay in Knowledge, asking a question, reading the answer, and moving on. So we made a bet: if we deliberately build onramps from Knowledge to Application to Reflection inside Coach's activities, more learners will take them.

How we measured KAR, and what we found

To turn KAR from a design philosophy into something we could check, we built evaluations specific to a sample of activities. The core is an LLM-as-judge setup: a classifier reads a full learner conversation and scores it for evidence of Knowledge, Application, and Reflection. We calibrated those classifiers against human raters, with the two agreeing around 85% of the time.

It's worth being precise about what "Application" meant to the classifier. A conversation counted as showing Application when the learner actually attempted the thing, writing a draft answer or a first message in their own words and then engaging with feedback on it, rather than only asking what to say. Asking "what should I put in my outreach message?" is Knowledge. Writing the message and reacting to a critique of it is Application. We scored for the second.

We ran this across about 350 anonymized learner conversations from a set of partners including PwC, LA Tech, Athens State University, and Merit America, comparing our original activities against redesigned versions built around explicit KAR onramps.

The clearest signal came from the mock interview and career exploration activities. In the original version of the mock interview activity, about 18% of conversations showed Application, with learners practicing their answers rather than asking what a good answer would be. In the redesigned version, that rose to 30%. Our Explore Career Pathways activity moved similarly, from 12% to 31%.

Reflection remained rare in both original and redesigned activities, at 1 to 3% of conversations. We're not sure yet whether that's a design problem we haven't solved or a sign of something more basic: that making sense of what you just tried may be the phase that depends most on another person. If it's the latter, prompting may not be enough on its own, and part of Coach's job could be connecting the learner with someone to do that thinking with.

Ultimately, these experiments are a small but positive signal that you can build real onramps from Knowledge to Application in an AI tool. Most people still arrive wanting an answer, the same as they do with any chatbot; the difference is how many go on to try something with it.

Still raising the bar

Coaching is hard for humans, and it's harder for AI. We don't think the answer is to lower the bar. It's to keep raising it, methodically, with frameworks grounded in research and tested against real learner outcomes. For now that means extending this measurement beyond the handful of activities we studied here, and treating what it surfaces as a target to refine each activity against.

If you’d like to learn more, we’d be glad to connect.

Next
Next

From Biology to Podiatry: How Ashley Found Her Footing with Coach