Research releases, technical notes, and what we're learning building domain-specific AI — published as we go, measured before we say it.
A pre-registered negative result: fine-tuning on survey microdata failed to transfer to simulated survey respondents.
We fine-tuned a ~9B open-weight model on 142,052 examples of Pew survey microdata, fixed the pass/fail criterion before training, and evaluated it raw against our five-model frontier ensemble under an identical harness. The adapter mastered its training objective (perplexity 1.75) — and failed the deployed task decisively: 46.92pp error versus the ensemble's 12.53pp, with a sixth of its respondents unable to answer at all. The first Track 01 report delivers on its promise: we said we'd publish what holds and what doesn't. This one didn't — and why it didn't is the finding. Updated Aug 28 with a post-publication validity audit: quantization exonerated, the reasoning-template mechanism characterized — the verdict unchanged.
An open specification for population parity, and a stress-test protocol for sycophancy.
The yardstick behind our numbers, published so it can be checked rather than trusted: a formal parity metric (per-item, aggregate, subgroup, and spread-preservation forms), a three-condition sycophancy stress test that measures how far a panel bends when the question leaks what the asker hopes to hear, fair-comparison rules any vendor can be run under, and the parity-card reporting format. Specification first, results next — measured where everyone can watch.
Census-grounded agent populations that disagree like real ones.
Our first technical report. How CrowdOS builds a persistent panel of 4,000+ OCEAN-modelled archetypes across 21 markets, measures parity against Pew, ANES, GSS, and OpinionQA, and instruments against sycophancy — the failure mode where synthetic panels collapse into agreeable consensus. With an honest frame around our 91.6% headline figure: what it means, what it doesn't, and where synthetic panels break.
Our parity cards under the open specification — per instrument, per market, across model providers.
CrowdOS run through parity/1.0 and SST/1.0 exactly as MM-TR-002 prescribes, with every card published whole — including the restated headline figure, whichever way it moves.
Working on adjacent problems, or want to be notified when we publish? Write to the research team.