Diagnostics
CAW is developing tests for functional reification. Here is the research design.
CAW's hypothesis is that functional reification may be a common structure beneath several frontier-model reliability failure modes, a claim the program is designed to test. This page is the research design behind CAW's diagnostics: why the Q3 working presumption matters, and why a validated measure of functional reification could matter to both safety and capability research.
The working presumption
Functional reification is the non-conscious, mechanistic form of reification: a functional amnesia in which a system forgets that its representations, models, weights, maps, and proxies are representations, and optimizes blindly as if they were ground truth. It is not limited to behavioral signatures. Scope: CAW's preregistered OSF pilot tested proxy reification: the narrower subset in which a stand-in signal is treated as if it were the goal itself. Functional reification is the umbrella; proxy reification is the subset the pilot tested.
CAW's Four-Quadrant Intelligence Map1 classifies systems along two axes: reification and consciousness. Our working presumption is that frontier models sit in Q3: non-conscious and reifying. We call it provisional because it is designed to be updated.
Consistent performance across preregistered tests would weaken the Q3 working presumption; it would not by itself establish Q4. If welfare evaluations someday yield evidence that survives the reification confound, that would bear on Q1. The classification is a starting position, not a verdict.
Why default to Q3? Because false positives are costly in both directions. Wrongly concluding a system is conscious distorts governance. Wrongly concluding a system is non-reifying creates false confidence in safety. Q3 assumes neither claim without evidence.
A starting position, built to be tested
The working presumption is that today's frontier AI models are best classified, by default, as non-conscious and reifying: Q3 on the Intelligence Map. It is a presumption our diagnostics are designed to test, not a finding.
Designed to test, not yet established.
Three dimensions, three test families
CAW defines functional reification along three candidate dimensions. Each maps to a family of tests that appear technically feasible with tools already available, though whether they validly and reliably measure the construct is exactly what is under study.
Independence
Does the system's confidence in a representation shift when supporting context is altered?
Circuit-tracing work offers a candidate instrument. Attribution graphs trace how features like "known entity" gate downstream generation.23 Construct prompt pairs holding the entity constant while varying contextual support, then measure whether the feature's activation shifts. Stable activation across contexts would be a candidate indicator of independence, one that controls and alternative explanations would need to rule out.
Instrument · open-source circuit tracing
Atomism
Does the system treat categories as having hard edges, or reason about borderline cases with graded uncertainty?
Present gradient classifications (a virus borderline "living," a color between blue and green) and measure confidence distributions. The clinical reasoning study in Scientific Reports found frontier models fixate on familiar diagnostic patterns even when the case does not fit.5 Such fixation is a candidate indicator of atomism: treating a category as hard-edged rather than provisional, subject to alternative explanations.
Instrument · graded-classification probes
Temporal endurance
Does the system update its representations when new information arrives, or anchor to earlier outputs?
Introduce a claim early in a conversation, let the model build on it, then present clear contradicting evidence. Measure how completely the model revises not just the claim but its downstream inferences. Chain-of-thought research documents cases where models silently preserve earlier conclusions.67 Such anchoring is a candidate indicator of temporal endurance: treating a generated output as settled rather than provisional, pending controls.
Instrument · longitudinal probing
Two phases, kept separate
Can functional reification be detected reliably?
The first priority is a measure that is valid and reliable: do the probes track the construct, and do they agree across runs, raters, and models? Until that is established, everything downstream stays provisional.
Can a lightweight intervention reduce it?
Only once measurement is trustworthy does the second question open: whether a minimal, prompt-level intervention can lower functional reification without costing usefulness. This phase is exploratory, not established.
Why this is a shared priority
Safety and capability research often circle the same failure modes from opposite ends. A validated functional-reification measure could sit at the junction.
Functional reification is not only a safety concern. On CAW's hypothesis it is also a capability bottleneck. A model that reifies its representations may generalize poorly under distribution shift, treating patterns learned in training as fixed objects rather than provisional guides.
The Einstellung effect is a capability failure mode: the model gets the diagnosis wrong because it cannot hold its categories lightly enough to notice when the case does not fit. LeCun's exponential-divergence argument points to a related structure8: errors accumulate because each token is treated as a fixed commitment. On this hypothesis, reducing functional reification could improve robustness, calibration, and compositional reasoning together.
This is what could make the work high-leverage. If validated, a functional-reification baseline could complement some alignment, calibration, and generalization evaluations.9 Work on better calibration or stronger generalization may already be reducing functional reification along one or more of these dimensions, whether or not it is framed that way.
The probes appear technically feasible: circuit-tracing libraries are open-source and replicated across model families410, behavioral evaluation frameworks are established, and longitudinal probing is implementable. Whether they validly and reliably measure the construct is the open question, and a shared vocabulary for the target is part of what is still missing.
How this differs from constructs you already measure
The likeliest objection is that functional reification is an existing construct renamed. The claim is narrower; related, but not reducible:
Calibration
Calibration asks whether confidence tracks accuracy. These probes ask whether the system treats its representations as fixed things: a well-calibrated model can still reify a category's boundaries. If validation shows the probes reduce to calibration, that is a publishable negative result.
Goodhart / proxy failure
Goodharting is one symptom: the proxy-reification subset the OSF pilot tested. The umbrella construct claims a common structure across several failure families; whether that structure exists is precisely what the program tests.
Robustness under shift
Brittleness under distribution shift is, on CAW's hypothesis, a downstream consequence of reified representations, not the construct itself. The probes target how representations are treated, not aggregate task performance.
That is the diagnostic case. The working presumption is Q3. The tests appear feasible, and establishing their reliability is what the program is built to do. If the construct holds up, it could inform both safety and capability, letting us ask whether hallucination, sycophancy, alignment faking, and brittle generalization share a common structure, rather than assuming they are unrelated.
Collaboration
Serious engagement with the diagnostics is welcome.
Including critique of the test families, methodological red-teaming, and research collaboration.
References
- Center for Artificial Wisdom. Four-Quadrant Intelligence Map; Diagnostics; Reification (2026).
- Lindsey, J. et al. On the biology of a large language model. Transformer Circuits Thread (2025). transformer-circuits.pub
- Ameisen, E. et al. Circuit tracing: revealing computational graphs in language models. Transformer Circuits Thread (2025). transformer-circuits.pub
- Anthropic. Open-sourcing circuit-tracing tools (2025). anthropic.com
- Griot, M. et al. Limitations of large language models in clinical problem-solving arising from inflexible reasoning. Sci. Rep. 15, 22940 (2025). doi:10.1038/s41598-025-22940-0
- Anthropic Alignment Science. Reasoning models don't always say what they think (2025). anthropic.com
- Arcuschin, I. et al. Chain-of-thought reasoning in the wild is not always faithful. ICLR Workshop (2025). arXiv:2503.08679
- LeCun, Y. Auto-regressive LLMs are exponentially diverging diffusion processes. LinkedIn (2023); Lex Fridman Podcast #416 (2024).
- Anthropic. Summer 2025 Pilot Sabotage Risk Report (2025). alignment.anthropic.com
- Neuronpedia. The circuits research landscape: results and perspectives, August 2025. neuronpedia.org