Trace Interactive explorer →
Research update · September 2026

Circuit-Tracing MedGemma

A cross-layer transcoder for a medical vision–language model: how it was built, how faithful it is, and what it caught.

Kevin Jin·Duke University

The method follows Anthropic’s Circuit Tracing. Full details accompany the release.

Summary

MedGemma is Google’s open vision–language model for medical images. We trained a 6.24B-parameter cross-layer transcoder (CLT) over all 34 of its layers: a sparse, interpretable replacement for the model’s MLPs that reads image tokens as well as text. To our knowledge it is the first CLT released for a medical VLM. Substituted for every MLP, it reproduces MedGemma’s next token on 60 of 60 held-out chest X-rays. With attention frozen it yields exact attribution graphs from image features to the answer, and it exposed a readout artifact that inflates feature-intervention effects about 150×. The weights are public, and the graphs can be browsed in the interactive explorer.

1What we built#

Much of a transformer’s computation happens in its MLP layers, but their activations are dense vectors that do not line up with human concepts. A transcoder is trained to predict each MLP’s output from its input through a large, sparse dictionary of features. Of the 2,048 features per layer in ours, only about ten fire on a given token, and each tends to track one thing: the lung base, a chest tube, the “AP PORTABLE” stamp. With the transcoder standing in for the MLPs, the model’s computation can be read feature by feature.

MLP activations dense, unlabeled transcoder Sparse features lung base tube AP stamp ~10 of 2,048 active, each interpretable
Figure 1. What a transcoder does. The same layer, re-expressed as a few active features out of a large dictionary (illustrative).

Ours is cross-layer: a feature read at one layer writes to every layer after it, so a single feature can account for computation the model spreads across depth. This is the substrate the attribution-graph method was built around. The CLT uses TopK sparsity (k = 32) with 2,048 features on each of MedGemma 1.5’s 34 layers, 6.24B parameters in all, trained from scratch on one H100 on activations streamed live from the model over 50,000 NIH chest X-rays and CheXpert Plus reports. As a single tensor the decoder would exceed 231 elements and overflow the int32 indexing in bitsandbytes’ 8-bit Adam, so it is stored as one tensor per write offset and decoded as a running sum, which also keeps activation memory flat.

INPUT WHAT WE READ Chest X-ray+ report MedGemma(34 layers) Cross-layertranscoder Attributiongraph run &capture replacethe MLPs freezeattention
Figure 2. The pipeline. Capture MedGemma’s activations on an X-ray and report, replace its MLPs with the CLT, then freeze attention to get an exact graph from features to the answer (§4).

2Images need their own dictionary#

MedGemma sees each X-ray as 256 image tokens placed ahead of the text. Their activations look nothing like text’s, and a transcoder trained the usual way, on text alone, does not transfer. Its fraction of variance unexplained (FVU, the share of the MLP output it fails to reconstruct) on image tokens is 1.13, worse than simply predicting the mean, and at layer 0 it is 354× worse than on text. Balancing image and text tokens in every training batch brings image FVU to 0.19 at no cost to text.

1.0 = guessing the average 0.191.13 0.200.19 Text only Text + images text tokens image tokens
Figure 3. Reconstruction error (FVU; lower is better) on each kind of token, by what the transcoder was trained on. A text-only dictionary cannot represent what the model does with the X-ray.

3How faithful it is#

The standard test is replacement: substitute the transcoder for all 34 MLPs, with no correction term, and check whether the model still behaves the same. On 60 held-out image/report pairs the replacement model predicts the original top-1 next token every time, with a mean KL divergence of 0.044 between the two output distributions.

60 / 60 same next token
Figure 4. Each dot is a held-out chest X-ray on which the replacement model matches the original model’s next token.

This holds on the prompt the CLT was trained with. On a constrained yes/no question the text side degrades (KL 13.5), while image-token reconstruction is identical across prompts: under causal masking the image tokens come before the question and cannot see it. Retraining on mixed prompts cuts the probe KL 13.5×, while doubling the dictionary does not help, so the bottleneck is prompt coverage, not capacity.


4Exact attribution graphs#

To explain one answer, we build a local replacement model for that X-ray: freeze the attention patterns and normalization at the values they take on this input, and add back the small per-position gap between the CLT and the true MLPs. What is left is linear in the features, so every edge in the resulting attribution graph is exact, and the edges into the answer sum to its logit (the model’s raw score for that token). We checked each property: the logit matches the original exactly; scaling features moves it along a straight line (deviation 0.094, versus 11.2 with attention left live, about 120× closer); and the edges sum to the logit within bf16 precision.

EARLY LAYERS ANSWER “Yes” image feature text feature
Figure 5. An attribution graph, simplified. Nodes are features; each edge is how much one feature contributes to the next, flowing toward the answer. Freezing attention makes these contributions exact and additive, at a cost: where the model looks is held fixed, so effects carried by the attention pattern itself are invisible. Step through real graphs in the explorer.

5What the features see#

Labeling features by their top-activating image patches returns concrete radiology: lung bases, the mediastinum, tubes and catheters, shoulders and spine, and acquisition cues such as the “AP PORTABLE” stamp on bedside films. In layer 30, 25 of the top 30 features receive a specific anatomical or acquisition label. The features also know where the evidence is. Occluding small patches of an X-ray reveals which regions actually drive the model’s answer, and a linear probe on the features picks out those regions with AUC 0.88, against 0.62 for raw pixels and 0.48 for a shuffled-label control.

Figure 6. Six features, each shown by the image patches that activate it most. Browse the full set in the explorer.

6A readout artifact#

The usual way to test a feature is to intervene on it and watch the output. On films where MedGemma falsely calls cardiomegaly (an enlarged heart), we negated the 80 image features most associated with the call. P(yes), read as the softmax probability of the “yes” token, fell from 0.92 to 0.63: an apparent handle on the finding.

It is not one. P(no) did not rise; the lost probability went to other tokens. The intervention made the model less likely to answer in yes/no form, not less convinced. Renormalized over yes and no, P(yes) stays at 1.00, a change of 0.002. The raw readout inflates the effect about 150×.

BeforeAfter yes 0.92yes 0.63 other tokens 0.37 other 0.08 01
Figure 7. Where the probability goes. Negating the features shrinks “yes” but hands the mass to unrelated tokens; “no” stays near zero throughout, so the model’s decision between yes and no is unchanged.

Intervening on the input, by contrast, does move the decision. Blacking out the single most causal patch drops the normalized P(yes) by 0.72, while blacking out a patch at the border drops it by 0.001. The call is driven by a specific image region, the features encode where that region is, and yet editing the features does not change the decision.

Black out the key patch0.718Negate the 80 features0.002Black out a border patch0.001
Figure 8. Change in the yes-versus-no decision under three interventions. Removing the evidence from the image works; removing it from the features does not.

The same kind of control caught one of our own earlier claims. We had localized a false cardiomegaly call on portable (AP) films to the attention pattern, because patching in a standard-film pattern removed it. Patching in the pattern from another portable film removes 87% as much, so the effect was mostly generic disruption rather than anything specific to film type. We have withdrawn it.


7Limits#

One model and one training run. The released CLT is faithful on its training prompt, not yet on arbitrary questions (§3). A per-layer transcoder trained on the same data reconstructs about as well at a seventeenth of the size, so the case for the cross-layer design rests on circuit tracing, not reconstruction. Feature labels are automatic and have not been validated by a radiologist.

References
  1. Lindsey, Templeton et al., Sparse Crosscoders for Cross-Layer Features and Model Diffing. Transformer Circuits Thread, 2024.
  2. Ameisen, Lindsey et al., Circuit Tracing: Revealing Computational Graphs in Language Models & On the Biology of a Large Language Model. Transformer Circuits Thread, 2025.
  3. Templeton et al., Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread, 2024.
  4. Transcoders in vision–language models. Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models, 2026 (preprint).  ·  Circuit Tracing in Vision-Language Models, 2026 (preprint): per-layer transcoders and attribution graphs on general Gemma-class VLMs.
  5. Google DeepMind & Google Research, MedGemma Technical Report, 2025.