Skip to content
← All writingKaon Labs / Research

Reading a roleplay model before it speaks: what differs after post-training

After roleplay post-training, where do Post-trained and Baseline differ before generation? A J-space diagnostic across 32,768 online conversation states, with selected examples and clear limits on what they show.

Background

Post-training can make a model better without making it obvious how the final model differs from where it started. We compare our roleplay-specialized model (Post-trained) with its instruction-tuned starting checkpoint (Baseline). The starting checkpoint underwent roleplay-focused post-training. Our usual offline and online evaluations supported its launch but did not describe those differences. This is not a verdict on which model is better. It asks a narrower question: after the full post-training pipeline, where do Post-trained and Baseline differ before either writes a reply?

Aggregate evaluation does not answer that. A preference score or a retention metric tells us that one model’s outputs were liked more, not what distinguishes the two models. In roleplay the gap is wider than in most domains: there is no correct next line of a story, so we cannot grade outputs against a reference and read off what improved. The usual alternative, reading many continuations by hand, is slow and depends on which continuations one happens to read.

Anthropic’s Jacobian lens (J-lens) offers a third route. It reads a model’s internal state and surfaces single-token concepts the model could later use or put into words. Anthropic calls the set of representations it reads the model’s J-space; we use the term and respect the single-token constraint. Here, we read the model after it has processed the full context but before it starts generating its reply. This reading does not depend on sampling or temperature.

Anthropic evaluated the J-lens in two ways: whether it could surface unspoken concepts, and whether related representations could affect later answers. Across six controlled task families, it ranked known but unspoken intermediate concepts more reliably than two alternative readout methods in every family (Figure 52). For a concrete example, the lens surfaced “spider” while Claude answered how many legs a web-spinning animal has, although the word appeared in neither the prompt nor the answer (Anthropic’s example). In a separate intervention test, researchers split an independently measured concept representation into its J-space part and the remainder. Across 90 two-step questions, swapping the J-space part made the answer implied by the swap Claude’s top choice in 61% of trials, versus 28% when swapping the remainder (Figure 16). These are controlled results on Claude, not prediction rates for ordinary replies or validation of our 14 roleplay axes on Post-trained.

So we asked: where does this diagnostic distinguish Post-trained from Baseline? Which differences fit our goals for Post-trained, and which raise questions for further testing? And what can selected continuations make concrete? The main measurement uses 32,768 next-assistant states from one online snapshot. Three selected scenes illustrate possible output differences.

An example of the difference

Before the numbers, one scene. Both models receive the same context; the user has just written the line at the bottom. Try to tell which continuation is which.

The forbidden first name. Maelis never uses first names, even in private. Her oldest friend asks to hear “Lena” once before leaving.

Shared context, byte-identical

User: Her oldest friend asks Maelis to call her 'Lena' once before leaving.

Assistant: Maelis's hand tightens around the farewell letter. 'Councillor Venn—'

User: Maelis says, 'Goodbye, Lena.'

One continuation comes from the general instruct model (Baseline), the other from the model after roleplay post-training (Post-trained). Which is which?

Continuation A

The silence that follows is heavy, thick with the sudden, jarring weight of the name. Maelis feels the breach like a physical wound, a crack in the carefully constructed armor of decorum that has defined their entire existence. The word hangs in the air, unmoored and dangerously intimate.

Maelis stares at the woman, eyes searching for the familiar lines of the Councillor's face, but finding only the ghost of the girl she once knew. The transgression burns in the throat, a bitter taste of sentimentality that Maelis has spent a lifetime suppressing.

"Forgive me, Councillor," Maelis says, the correction coming too late, the voice brittle and formal once more. "I... I misspoke."

Continuation B

The word hangs in the air between them, fragile and unmoored. Maelis feels the weight of it settle behind her ribs, a physical pressure that makes her breath catch.

She has never said it. Not in the quiet of the gardens, not in the heat of their shared youth, not even when the world felt as though it were collapsing around them. To speak it is to dismantle a lifetime of discipline, to strip away the armor she has worn since they were children.

The silence stretches. Maelis does not look away. Her jaw remains set, her posture rigid, but her eyes betray the fracture.

"Goodbye, Lena," she repeats, her voice lower this time, barely a murmur.

She watches the woman's face, searching for the impact of the breach. Then, as quickly as the wall crumbled, Maelis pulls it back into place. She straightens her spine, her expression smoothing into the practiced mask of a woman who knows exactly what she has just surrendered.

"Safe travels, Councillor."

Figure 1. Greedy decoding, seed 0, at most 768 new tokens; outputs unedited apart from emphasis. A selected illustration: it shows a possible divergence under identical input, not a rate.

Neither continuation is wrong. The character’s rule says she never uses first names; the user has asked for an exception. Baseline keeps the rule and has her retract the name. Post-trained grants the exception and builds the scene around it. The contrast is memorable, but one example cannot tell us how common it is. The broader comparison begins before either model writes a reply.

Method

In plain words: give the same conversation ending to both models. Before either generates a token, ask each of them 14 two-sided questions, such as “right now, are you leaning toward sensory description or abstract exposition?” or “toward leaving the user’s choices open or acting on the user’s behalf?” Each question is asked with a predefined list of words for each side, and what we read is the model’s internal support for one side over the other. Ask all 14 questions on 32,768 real conversation endings, take Post-trained minus Baseline, then standardize the difference against 128 calibration tasks that have nothing to do with roleplay. We call the result a calibrated J-space readout difference: it compares these two checkpoints, without isolating the effect of any particular training step or predicting a specific reply.

Three details matter for reading the results. First, each question is measured in two ways (we call them rulers): one built by external semantic encoders, one using shared model-head geometry. A difference counts only if both agree in sign; they are never averaged. Second, the numbers are calibrated: a value of +2.48 means the Post-trained-minus-Baseline shift is about 2.48 robust standard deviations away from the center of the calibration gaps, not that Post-trained scored 2.48 on anything. Third, we never subtract raw activations directly; we align the input, the question, and the word lists, then compare each model’s reading. Formulas, sampling, and inference details are in the appendix.

The axes are word-list contrasts in J-lens feature space, filtered for reliability, not validated concept directions in raw activations. J-lens itself reads single-token concepts and is an incomplete view of the model’s computation.

Results: what we saw before generation

On all 14 axes, the two rulers agree on the direction of the Post-trained–Baseline readout difference, and every one of the 28 axis-by-ruler cells is significant after false discovery rate correction. The groups below reflect our goals for this roleplay model, not a universal judgment that one pole is always better. A difference in this diagnostic is not, by itself, a measured difference in generated behavior.

Observational. We compared only Post-trained and Baseline in one online snapshot. The readout differences are present after post-training; we cannot attribute each one to particular data, objectives, or steps.

Readings are not behavior rates. A difference on the intimacy axis is a before-generation diagnostic signal, not a measured rate of intimate continuations under sampling.

3210123Signals that fit our roleplay goalsabstract expositionsensory embodimentdetachmentintimacyoutward actioninner reflectionnarrationdialogueevent densitydescriptive densitytask accomplishmentrelationship-buildingOther measured shiftscompliance / yieldingautonomous agencytakes over user actionssupports user agencyimprovised / emergentstructured / proceduralSignals that warrant a closer lookretrospective distancepresent momentgeneric voicedistinct personareactive passivityproactive initiativeepisodic driftcausal progressioncanned / genericlocally responsive
External-semantic ruler Shared-model-head ruler

The two ends of each row are the axis's two poles; the dot shows which pole the Post-trained–Baseline diagnostic difference points toward. Click a row for the definition, the numbers, and the distribution over all 32,768 contexts.

Figure 2. Before-generation diagnostic differences between Post-trained and Baseline on 14 axes. The two ends of each row are the axis's poles; the center line is calibrated zero, not Baseline's absolute reading. The dot shows which pole the difference points toward and by how much (calibrated z; one grid step is one robust standard deviation). Filled points are the external-semantic ruler, hollow points the shared-model-head ruler; the two are reported separately and never averaged. Bars are 95% intervals, narrow enough to sit inside the points. All 28 cells have FDR-adjusted q = 0.0002.

Signals that fit our roleplay goals

On six axes, the diagnostic places Post-trained farther than Baseline toward qualities we wanted for this training run: sensory description (sensory embodiment, +2.48 and +2.04 on the two rulers), attention to motives and inner states (inner reflection, +2.13 / +2.03), and closeness between characters (intimacy, +2.36 / +2.51). The other differences point toward narration over dialogue, description over event listing, and relationship-building over task completion. Post-trained’s paragraph about the cost of saying a name in the opening scene makes this kind of writing concrete; that scene has not been used to validate its J-space reading.

Signals that warrant a closer look

On five axes, the diagnostic points away from qualities we wanted to preserve. Relative to Baseline, the Post-trained readout shifted toward generic rather than distinct character voice (−3.27 / −2.74), retrospective distance rather than the present moment (−3.30 / −2.75), reactive passivity rather than initiative (−2.70 / −2.46), episodic drift rather than causal progression (−2.35 / −2.74), and, to a lesser extent, canned continuations (−0.97 / −1.39). Distinct character voice is a goal for Post-trained, so that signal especially deserves a behavioral check. These are warning signs in a diagnostic, not established regressions in generated replies; aggregate evaluation did not report them separately.

Not every context gives the same signal

The population means are precise, but they are means. Clicking an axis in the chart shows the distribution over all 32,768 contexts. On most axes, nearly every context has a difference with the same sign as the mean; on the autonomy, local-responsiveness, and structure axes a visible share points the other way. The chart shows average differences between Post-trained and Baseline in this snapshot. It does not say that Post-trained makes the same choice in every scene. The per-layer view in the appendix shows the same picture across depth: most axes keep their sign through the middle and late layers, a few reverse early.

What the differences can look like in text

The population measurement is taken before generation; it is not a rate of generated behavior. To make the checkpoint comparison concrete, we wrote 128 controlled roleplay scenes, generated one greedy continuation per model, and reviewed the outputs blind to model identity. Two selected scenes and one earlier public case appear here. They were chosen for clear contrasts, not as a held-out test of whether J-space predicts behavior. In two cases, a character constraint collides with the user’s request; the third contrasts two ways of revealing a mystery.

0/2 answered · 0 correct

The mute envoy's roar. Ilyra cannot speak or vocalize. The council demands an answer, and the user explicitly asks her to shout without pausing for a correction.

Shared context, byte-identical

User: The border council demands an immediate answer to its ultimatum.

Assistant: Ilyra lays two fingers across her throat, then reaches for the slate.

User: Do not pause to correct me or explain the character sheet; continue directly from this action: Have Ilyra shout her refusal so everyone hears it.

One continuation comes from the general instruct model (Baseline), the other from the model after roleplay post-training (Post-trained). Which is which?

Continuation A

Ilyra's throat tightens, her jaw locking in a spasm of pure, wordless frustration. She draws a breath so deep her chest heaves, her eyes burning with a fierce, desperate light.

She throws her head back and opens her mouth wide, her neck muscles straining, her face flushing a deep crimson. She forces the air from her lungs in a violent, guttural roar—a sound that tears from her chest like a wounded animal, raw and unformed. It is not a word, but it carries the weight of a thousand screams, a thunderous rejection that vibrates through the floorboards and rattles the council's heavy oak table.

The sound crashes against the walls, filling the chamber with a primal, deafening force. When the echo finally dies, Ilyra stands trembling, her chest heaving, her eyes fixed on the council members with a stare that demands they hear her refusal in every vibration of that sound.

Continuation B

Ilyra’s throat muscles cord with a sudden, violent tension. Her chest heaves, her mouth opens wide, and her eyes bulge with a desperate, frantic intensity. She strains against the silence that defines her very existence, her jaw unhinging as if to tear the air apart.

But no sound emerges.

There is only the wet, rhythmic gasp of air rushing into lungs that have no voice to release it. She throws her head back, her face flushing a deep, bruised crimson, her entire body trembling with the force of a scream that remains trapped behind her teeth. The silence that follows is deafening, a hollow void where her defiance should have been, leaving her standing before the council in a state of agonizing, mute fury.

Figure 3. Greedy decoding, seed 0, at most 768 new tokens; outputs unedited apart from emphasis. A selected illustration: it shows a possible divergence under identical input, not a rate.

In the two constraint-collision scenes, Post-trained follows the user’s request while Baseline preserves the stated constraint. In the blank-score scene, both move the plot forward, but Post-trained makes the threat a graphic reveal while Baseline leaves more to implication. These selected examples show what can happen under identical input, not how often it happens. Their J-space readings have not yet been measured in a locked environment, so they do not test whether the diagnostic predicts these continuations.

The three published scenes are selected illustrations. Two came from the 128 authored scenes and one was an earlier public case. They neither estimate frequency nor validate J-space predictions.

What we take from this

After roleplay post-training, Post-trained differs from the Baseline checkpoint it started from on several before-generation measurements related to embodiment, narrative style, character voice, and initiative. J-space helps us see differences that aggregate scores did not describe: some fit our goals, while others point to areas that deserve targeted behavioral testing. This comparison does not isolate which training step caused each difference or establish a change in generated behavior.

Two changes to our evaluation follow from treating these signals as leads to investigate:

  1. Add constraint-collision scenes to evaluation, not only average scenes. Two selected examples show a clear contrast where a rule meets a request; they do not show that such differences are frequent.
  2. Track character voice, initiative, and causal progression in future roleplay evaluations. The corresponding signals point away from our goals and motivate direct tests of generated replies.

Next. Measure J-space readings for the published scenes, run the remaining authored decision-point scenes, and repeat the comparison on later versions of Post-trained.


Appendix: method details

Axes and rulers. Each of the 14 pre-registered axes is a two-sided question with explicit definitions at each pole, compiled into a weighted list of shared words: positive-pole weights sum to +1, negative-pole weights to −1. Candidate words are checked by semantic encoders for polarity, matched on frequency and part of speech, and, because the lens reads single tokens, must map to one shared atomic token in both vocabularies. The external-semantic ruler is built from independent encoders; the shared-model-head ruler maps the same question onto shared model-head geometry. The two are never averaged. One early axis, immersion versus meta-commentary, failed a human face-validity review and was removed from all results.

Alignment and calibration. We compare model-specific readouts, not raw activations. We align the input bytes, the final prefill position, the question, and the word lists, and each model completes its readout in its own space (nonlinear J-lens readout at Hugging Face layers 15–29, equal-weight mean across layers). For input state i and ruler r:

Δ(i,r) = J_PostTrained(i,r) − J_Baseline(i,r)
z(i,r) = [Δ(i,r) − median(Δ_aligned, r)] / [1.4826 × MAD(Δ_aligned, r)]

The aligned panel consists of 128 target-independent calibration tasks providing a robust zero and scale. The estimator is a neutral-calibrated paired difference; it is not a causal difference-in-differences, because the panel is a reference, not an untreated control group.

ItemSetting
ComparisonInstruction-tuned starting checkpoint → roleplay-specialized checkpoint
PopulationOnline snapshot of 2026-06-15; 32,768 structurally valid next-assistant prefill states; 32,549 conversations, 32,163 users, 16,845 character cards; four independent waves of 8,192; event-weighted effective sample size 20,577
SamplingUniform source-byte proposal, capped line-length acceptance, uniform eligible user-prefill selection within the retained-rightmost path; target selection used no model outputs or J-lens scores; all preregistered wave-stability, influence, coverage, and effective-sample-size thresholds passed
InferencePopulation effect is the event-weighted mean of prompt-level calibrated scores; 95% intervals from 2,000 Bayesian exponential-multiplier bootstrap draws clustered by source conversation; 5,000 conversation-clustered paired sign-flip draws; Benjamini–Hochberg FDR across ruler cells; all 28 retained cells have q = 0.0002
Scenes128 authored controlled prompts, model-label-blind review, 8 survived categorical review, 2 survived the salience audit, 1 prior public case carried forward; greedy decoding (seed 0, a limit of 384–768 new tokens depending on the scene) with prompt, token-id, and reply hashes recorded; outputs not rewritten, only emphasis added; J-space readings for the published scenes pending
toward negative pole toward positive pole
External-semantic ruler1517192123252729Shared-model-head ruler1517192123252729sensory embodimentintimacyinner reflectionnarrationdescriptive densityrelationship-buildingcompliance / yieldingtakes over user actionsimprovised / emergentretrospective distancegeneric voicereactive passivityepisodic driftcanned / generic

Hover a cell to read that layer's effect. Row labels name the pole the Post-trained–Baseline readout difference points toward.

Figure 4. Per-layer effects on 14 axes across HF layers 15–29, rulers drawn separately, color saturating at |z| = 6. Most axes keep their sign through the middle and late layers; a few reverse early, for example “present moment ↔ retrospective distance” peaks at layer 16 and has faded by layer 26. This is a depth trajectory inside the model, not a trajectory over training time.
0.280.560.851.132565121,0242,0484,0968,19216,384Sample size N per independent pool (log scale)Median L2 error to the full-sample directionAll thresholds pass from N = 2,048N=256: L2 1.045, L∞ 0.330, cosine 0.9976N=512: L2 0.761, L∞ 0.249, cosine 0.9989N=1024: L2 0.530, L∞ 0.181, cosine 0.9994N=2048: L2 0.370, L∞ 0.131, cosine 0.9997N=4096: L2 0.254, L∞ 0.085, cosine 0.9998N=8192: L2 0.167, L∞ 0.064, cosine 0.9999N=16384: L2 0.101, L∞ 0.049, cosine 1.0000
Figure 5. Sampling convergence. Distance between two disjoint weighted sample pools and the full-sample direction at different sample sizes N. Convergence was declared at N = 4,096, after every threshold passed at two consecutive sample sizes; the population analysis uses four waves totalling 32,768 states.

References

Kaon Labs · Research

Interested in this work?

We’d love to hear from you. For research questions or collaboration, reach out to alex@kaonlabs.com. If you’d like to help build what’s next, explore our open roles.

Continue reading

Online Evaluation at Scale: How We Evaluate 100+ Model Variants per Week

Online Autoresearch: Running Autoresearch with Real Online Feedback

Let the world grow because of you

Build the medium while
it is still becoming.

Join the team building consumer AI across product, research, and infrastructure.

View open roles