Blog

Notes from the lab.

Research releases, technical notes, and what we're learning building domain-specific AI — published as we go, measured before we say it.

August 2026 · Research · Negative Result MM-TR-003

Distribution Learning Is Not Respondent Simulation.

A pre-registered negative result: fine-tuning on survey microdata failed to transfer to simulated survey respondents.

We fine-tuned a ~9B open-weight model on 142,052 examples of Pew survey microdata, fixed the pass/fail criterion before training, and evaluated it raw against our five-model frontier ensemble under an identical harness. The adapter mastered its training objective (perplexity 1.75) — and failed the deployed task decisively: 46.92pp error versus the ensemble's 12.53pp, with a sixth of its respondents unable to answer at all. The first Track 01 report delivers on its promise: we said we'd publish what holds and what doesn't. This one didn't — and why it didn't is the finding. Updated Aug 28 with a post-publication validity audit: quantization exonerated, the reasoning-template mechanism characterized — the verdict unchanged.

Read the report → Download PDF
August 2026 · Research · Evaluation MM-TR-002

Measuring Synthetic Society.

An open specification for population parity, and a stress-test protocol for sycophancy.

The yardstick behind our numbers, published so it can be checked rather than trusted: a formal parity metric (per-item, aggregate, subgroup, and spread-preservation forms), a three-condition sycophancy stress test that measures how far a panel bends when the question leaks what the asker hopes to hear, fair-comparison rules any vendor can be run under, and the parity-card reporting format. Specification first, results next — measured where everyone can watch.

Read the report → Download PDF
July 2026 · Research · CrowdOS MM-TR-001

Calibrating Synthetic Society.

Census-grounded agent populations that disagree like real ones.

Our first technical report. How CrowdOS builds a persistent panel of 4,000+ OCEAN-modelled archetypes across 21 markets, measures parity against Pew, ANES, GSS, and OpinionQA, and instruments against sycophancy — the failure mode where synthetic panels collapse into agreeable consensus. With an honest frame around our 91.6% headline figure: what it means, what it doesn't, and where synthetic panels break.

Read the report → Download PDF
Coming Next In Preparation

The results, restated in the open.

Our parity cards under the open specification — per instrument, per market, across model providers.

CrowdOS run through parity/1.0 and SST/1.0 exactly as MM-TR-002 prescribes, with every card published whole — including the restated headline figure, whichever way it moves.

Working on adjacent problems, or want to be notified when we publish? Write to the research team.