Our products are applied research. Beneath each one is a set of questions the field hasn't answered — and a discipline we bring to answering them.
The largest frontier models are extraordinary generalists. But a generalist is not a specialist, and most valuable work happens inside domains with structure that general benchmarks never test.
Motif Motion is organised around a single wager: that the next decade of useful AI comes not from larger general models, but from making AI genuinely competent inside specific domains — with the field's own logic, constraints, and standards of evidence built in.
That work has three layers. We build systems around models today — retrieval, orchestration, versioned domain engines, and evaluation. We are increasingly building models of our own, fine-tuning and training weights on domain corpora. And underneath both, we build evaluation — because a domain-specific system that can't be measured against ground truth is just a confident guess.
The pages below describe the questions we're actively working on. Several are unsolved in the open literature; all of them are load-bearing for products people already use.
Can a small model that knows one domain deeply outperform a frontier generalist on that domain's home turf?
Today our products encode domain knowledge in the harness — in retrieval, engines, and prompting around a general model. The frontier we're moving toward encodes that knowledge in the weights themselves: fine-tuning and training our own models on domain-specific corpora so the competence is native, not retrieved.
We're beginning where our domain data is richest and most structured — Chinese metaphysics and population simulation — with the aim of models that are smaller, faster, and more faithful than a general model prompted to imitate a specialist. This is early-stage and honest work: we'll report what holds and what doesn't.
How closely can a population of AI agents track the real distribution of human belief — and where does it break?
CrowdOS is our testbed. We validate synthetic responses against national survey instruments — ANES, GSS, and Pew — and study the conditions under which a synthetic panel drifts from the real population it's meant to represent.
A central problem is sycophancy: the tendency of language-model agents to converge on agreeable consensus rather than represent the genuine spread and friction of real opinion. We instrument against it directly, treating disagreement as a signal to preserve rather than an error to smooth away.
How do we build systems that show their basis — and decline when the basis isn't there?
A recommendation you can't trace is a liability in every domain we work in. An investment surfaced by Bacon must carry the signals that justify it. An outfit proposed by Riff must be constrained to garments the user actually owns. A reading from Daymaster must derive from the computed chart, not generic text.
The shared research problem is grounding: reasoning architectures that attach an evidence chain to every output, and that refuse gracefully when the ground truth is absent — the opposite of confident hallucination.
Can centuries-old symbolic systems be rendered as precise, testable computation — and narrated faithfully by a model?
Bazi, Zi Wei Dou Shu, and Tarot are old symbolic systems with exact internal logic. Daymaster is, to our knowledge, the first engine to fuse Bazi and Zi Wei Dou Shu into a single coherent reading, and the first to integrate a Tarot dimension alongside them — bridging Eastern and Western traditions in one computed synthesis.
The research is twofold: formalising each system into a versioned, testable engine computed from true solar time, and then studying how a language model can narrate their synthesis faithfully — without collapsing back into generic horoscope text. It's a clean case study in grounding a model to a formal source of truth.
There is no single metric for "good domain AI." Each field supplies its own ground truth, and our evaluation is built to match it:
Where we can, we state limits as clearly as results. A synthetic population is a model of a population, not the population. A metaphysics engine is a faithful computation of a tradition, not a claim about fate. Saying so plainly is part of the method.
A pre-registered negative result on fine-tuning for simulated survey respondents.
A ~9B model fine-tuned on 142,052 examples of Pew microdata mastered its training objective (perplexity 1.75) — and failed the deployed respondent-simulation task decisively against a prompted five-model frontier ensemble: 46.92pp vs 12.53pp raw error under a pre-registered criterion. The first published report from Track 01, and the reason we hold "we trained our own model" claims — ours included — to same-harness, pre-registered comparisons. Updated Aug 28 with a post-publication validity audit.
An open specification for population parity, and a stress-test protocol for sycophancy.
The formal parity metric behind our calibration claims — per-item, aggregate, subgroup, and spread-preservation forms — plus a three-condition sycophancy stress test, fair-comparison rules for running any synthetic-panel system through the same procedure, and the parity-card reporting format. Published before our own results under it, deliberately.
Census-grounded agent populations that disagree like real ones.
How CrowdOS builds a persistent panel of 4,000+ OCEAN-modelled archetypes across 21 markets, measures parity against Pew, ANES, GSS, and OpinionQA, and instruments against sycophancy — with an honest frame around our 91.6% headline figure: what it means, what it doesn't, and where synthetic panels break.
Next in the series: our parity cards under the open specification — per instrument, per market, across providers. All notes are published on the blog; if you're a researcher on adjacent problems, we'd be glad to hear from you.