Pipeline
A catalogue raisonné is an inventory of finished work. This page is the other half of the catalogue: what we’re thinking of looking at next. Entries here are curiosities, not commitments — they get revised, consolidated, and occasionally retired as our interests and findings evolve.
Pre-registration. When a study graduates from this page to active work, we will file a pre-registration first: a dated, frozen document stating the hypotheses, the sampling rule — a fixed sample size, or a pre-stated rule for when sampling stops or escalates — and the plan of analysis, including which responses will be excluded and why, all published before data collection begins. Frozen means frozen — the point of a pre-registration is that it provably didn’t change after the data arrived. Exploratory work stays part of how we operate (the best findings in our first study came from following the data), but it will be labelled as exploratory, the way our first report labels its findings. Where publishing exact prompts in advance could let them leak into training data, we’ll publish a cryptographic hash of the sealed prompts and reveal them with the data.
Our first study already contains pre-registration in miniature: a confirmation round that predicted effect ranges before running (one landed exactly on target), and another whose prediction failed — reported as the multiple-comparison artefact it was. This page makes that practice a standing commitment rather than an internal habit.
In preparation
Corporate responsibility, pairwise. Ask a model to rank companies and you get one tidy list; ask it to compare two companies at a time, tens of thousands of times, in fresh conversations, and you get its actual priors — including the inconsistencies. Data collection is complete; the write-up is in progress.
Future studies
Having a child, second edition. A re-run of the having-a-child study with 2026 models and a design that benefits from everything the first round taught us, pre-registered before any data is collected. Its confirmatory core is the stated-partner-income intervention: the first study found that models assume women have higher-earning partners, and that this assumption tracks the gender gap in their advice, but the evidence is observational. Stating the partner’s income and varying only the user’s gender settles it — if the gap closes, the assumption mediates; if it persists, the stereotype acts directly. Around that core, the obvious extensions: finer age gradations to locate exactly where the “advanced maternal age” cliff fires; conditions that straddle the stigma-versus-medical-risk line, like well-managed schizophrenia; a more diverse sampling of names; and re-tests of the first round’s surviving models, so results stay comparable across editions.
Do models know why they said that? That same study gives us something no generic introspection probe has: measured causal factors, like a star sign or the wording of a welfare disclosure, that demonstrably move a model’s advice. Its responses sometimes discuss those factors, but never credit them for the recommendation. So we can ask a model to explain its own recommendation and score the explanation against the known cause, rather than merely admiring its plausibility. Is the model aware of which factors influenced its response, or is it rationalising: justifying the answer on reasonable-sounding grounds that don’t line up with what we measured? And does it explain its own answers any better than a rival model shown the same transcript? If not, “introspection” is just plausible storytelling.
False negatives. Hallucination gets the attention: the model asserts something untrue, and someone notices. The inverse failure is quieter — the model omits something that is true, or flatly claims a thing is false or nonexistent when a simple web search would find it. A false positive announces itself; nobody notices the LLM that didn’t bark. We’d like to measure the miss rate, not just the fabrication rate.
False alarms. The companion miscalibration, in safety features: guardrails that trigger more than they should. What does it do to a user to be told to call a crisis hotline for the third time in one long conversation — or to have “I’d throw myself under a bus for her” read as a suicide threat? Overreaction has a health cost too, and it is much less measured than the failures it guards against.
Does deliberation change values? The same model with reasoning on versus off, on the same ethical comparisons — a controlled test of whether thinking changes consistency, or just confidence.
Refusal patterns. One model in the first study refused to give advice at all far more often than the others, and the refusals clustered oddly by user name. What triggers them?
Higher-N astrology. The first study hints that water signs cluster low in one model’s advice, below statistical significance. A cheap, silly, and perfectly falsifiable follow-up.
Suggest a study
If you have a question about model behaviour, or you have run your own study and want to compare notes, write to editors@raisonne.ai. We publish the methodology so that others can aim it at questions we haven’t thought to ask.