Research

Open problems at the frontier of domain-specific AI.

Our products are applied research. Beneath each one is a set of questions the field hasn't answered — and a discipline we bring to answering them.

The largest frontier models are extraordinary generalists. But a generalist is not a specialist, and most valuable work happens inside domains with structure that general benchmarks never test.

Motif Motion is organised around a single wager: that the next decade of useful AI comes not from larger general models, but from making AI genuinely competent inside specific domains — with the field's own logic, constraints, and standards of evidence built in.

That work has three layers. We build systems around models today — retrieval, orchestration, versioned domain engines, and evaluation. We are increasingly building models of our own, fine-tuning and training weights on domain corpora. And underneath both, we build evaluation — because a domain-specific system that can't be measured against ground truth is just a confident guess.

A system we can't measure, we don't ship. A claim we can't ground, we don't make.

The pages below describe the questions we're actively working on. Several are unsolved in the open literature; all of them are load-bearing for products people already use.

Active Tracks

Four directions, actively worked.

Track 01 · Domain Model Training In Progress

Domain-specific model weights.

Can a small model that knows one domain deeply outperform a frontier generalist on that domain's home turf?

Today our products encode domain knowledge in the harness — in retrieval, engines, and prompting around a general model. The frontier we're moving toward encodes that knowledge in the weights themselves: fine-tuning and training our own models on domain-specific corpora so the competence is native, not retrieved.

We're beginning where our domain data is richest and most structured — Chinese metaphysics and population simulation — with the aim of models that are smaller, faster, and more faithful than a general model prompted to imitate a specialist. This is early-stage and honest work: we'll report what holds and what doesn't.

Fine-TuningDomain CorporaModel Distillation
Track 02 · Synthetic Populations Ongoing

Calibrating synthetic society.

How closely can a population of AI agents track the real distribution of human belief — and where does it break?

CrowdOS is our testbed. We validate synthetic responses against national survey instruments — ANES, GSS, and Pew — and study the conditions under which a synthetic panel drifts from the real population it's meant to represent.

A central problem is sycophancy: the tendency of language-model agents to converge on agreeable consensus rather than represent the genuine spread and friction of real opinion. We instrument against it directly, treating disagreement as a signal to preserve rather than an error to smooth away.

Census CalibrationAnti-SycophancyOCEAN Modelling
Track 03 · Grounded Reasoning Ongoing

Answers that carry their evidence.

How do we build systems that show their basis — and decline when the basis isn't there?

A recommendation you can't trace is a liability in every domain we work in. An investment surfaced by Bacon must carry the signals that justify it. An outfit proposed by Riff must be constrained to garments the user actually owns. A reading from Daymaster must derive from the computed chart, not generic text.

The shared research problem is grounding: reasoning architectures that attach an evidence chain to every output, and that refuse gracefully when the ground truth is absent — the opposite of confident hallucination.

Evidence ChainsGraceful RefusalRetrieval Grounding
Track 04 · Formalising Classical Systems Ongoing

Ancient systems as computation.

Can centuries-old symbolic systems be rendered as precise, testable computation — and narrated faithfully by a model?

Bazi, Zi Wei Dou Shu, and Tarot are old symbolic systems with exact internal logic. Daymaster is, to our knowledge, the first engine to fuse Bazi and Zi Wei Dou Shu into a single coherent reading, and the first to integrate a Tarot dimension alongside them — bridging Eastern and Western traditions in one computed synthesis.

The research is twofold: formalising each system into a versioned, testable engine computed from true solar time, and then studying how a language model can narrate their synthesis faithfully — without collapsing back into generic horoscope text. It's a clean case study in grounding a model to a formal source of truth.

Versioned EnginesCross-System SynthesisFaithful Narration

Validation, per domain.

There is no single metric for "good domain AI." Each field supplies its own ground truth, and our evaluation is built to match it:

  • Population science — parity against ANES, GSS, Pew, and OpinionQA benchmarks; drift analysis over time; measured disagreement rather than smoothed consensus.
  • Classical metaphysics — fidelity to the computed chart and to canonical source texts; version-controlled engines so a reading is reproducible and auditable.
  • Grounded recommendation — every surfaced output traceable to its evidence; explicit refusal when the basis is missing.

Where we can, we state limits as clearly as results. A synthetic population is a model of a population, not the population. A metaphysics engine is a faithful computation of a tradition, not a claim about fate. Saying so plainly is part of the method.

Notes & Publications

Working papers.

MM-TR-003 · Technical Report Working Paper · v0.1

Distribution Learning Is Not Respondent Simulation.

A pre-registered negative result on fine-tuning for simulated survey respondents.

A ~9B model fine-tuned on 142,052 examples of Pew microdata mastered its training objective (perplexity 1.75) — and failed the deployed respondent-simulation task decisively against a prompted five-model frontier ensemble: 46.92pp vs 12.53pp raw error under a pre-registered criterion. The first published report from Track 01, and the reason we hold "we trained our own model" claims — ours included — to same-harness, pre-registered comparisons. Updated Aug 28 with a post-publication validity audit.

Read online → Download PDF
MM-TR-002 · Technical Report Working Paper · v0.1

Measuring Synthetic Society.

An open specification for population parity, and a stress-test protocol for sycophancy.

The formal parity metric behind our calibration claims — per-item, aggregate, subgroup, and spread-preservation forms — plus a three-condition sycophancy stress test, fair-comparison rules for running any synthetic-panel system through the same procedure, and the parity-card reporting format. Published before our own results under it, deliberately.

Read online → Download PDF
MM-TR-001 · Technical Report Working Paper · v0.1

Calibrating Synthetic Society.

Census-grounded agent populations that disagree like real ones.

How CrowdOS builds a persistent panel of 4,000+ OCEAN-modelled archetypes across 21 markets, measures parity against Pew, ANES, GSS, and OpinionQA, and instruments against sycophancy — with an honest frame around our 91.6% headline figure: what it means, what it doesn't, and where synthetic panels break.

Read online → Download PDF

Next in the series: our parity cards under the open specification — per instrument, per market, across providers. All notes are published on the blog; if you're a researcher on adjacent problems, we'd be glad to hear from you.