Papers
An Event-Dated Corpus of US Trademark Prosecution and a Two-Sided Measure of Vocabulary Position
Patents count inventions, but most firms never patent; the goods-and-services language of trademark filings records what millions of firms tried to sell, dated to the day. This paper re-parses all 13.99 million US trademark case files into 242 million dated events, which turns the sworn five-year proof of continued use into a product-level survival outcome. On that corpus it builds two measures of each filing’s language: atypicality — how unusual the wording is for its industry — and lead — whether the wording ran ahead of or behind where the industry’s language was moving. The corpus, the measures and the code are public.
When Does It Pay to Be Early? Evidence from Every US Trademark Filing
Strategy offers two opposite answers to whether it pays to bring a new kind of offering to market first: pioneers keep the markets they open, or most pioneers fail and followers take what they proved. Using the lead measure on every registration from 2002 to 2018, the paper finds that being early usually costs: leading descriptions are cancelled more often at the five-year proof, and leading firms are no more likely to raise money, list or be bought. The cost is concentrated where an industry’s filing volume is growing fast, in platform and network businesses, and among filers without counsel; in a few industries, such as beer, leading filings survive more often.
An Assumption-Transparent Simulation of Automation and the Labor Market
Working paper, 2026 · the modular leg of the “Where Do the People Go?” program · not yet circulated.
Forecasts of AI’s effect on work usually stop at exposure scores; they do not say where displaced workers end up or what happens to wages. This is a labor-market simulation built from O*NET primitives — workers with graded capabilities, tasks with level requirements, jobs as bundles of tasks — calibrated so the pre-shock equilibrium reproduces BLS employment and wages by construction. Every assumption is a named, swappable module, so the question becomes which assumptions drive the answer rather than whose forecast to trust.
Generating a Credible Synthetic Workforce: Skill Endowments from Observed Occupational Transitions
Working paper, 2026 · the population-generation leg of the same program · not yet circulated.
Any simulation of who can move to which job is only as good as its synthetic workers, and public data never observe an individual’s full skill vector. This generator builds one by requiring that every worker be qualified for the job they actually hold, then grows skills along career paths observed in longitudinal surveys, and checks the population against employment rates, age and education mixes, and churn. The hypothesis is that qualification-by-construction plus observed transitions recovers a workforce realistic enough to forecast on.
How Soon Is Not How Long: Separating Care Urgency from Housing Permanence in Disability Waiver Triage
Working paper, 2026 · under internal review.
Placing someone on a lifetime disability waiver is close to irreversible — a present-value commitment of several million dollars that almost never moves to cheaper care — yet the entry pipeline only asks how soon help is needed, never how long it will be needed. Twenty years of person-level records from one county show the two questions come apart: a measurable share of urgent need is temporary, and the reasons written on the intake form predict persistence better than the urgency score does. The proposal is to separate the triage of speed from the decision of permanence.
The Failure of Medicaid Effectiveness Measures
In preparation, 2026 · why the measures a Medicaid system reports on itself do not track whether anyone got better.
Medicaid behavioral-health systems judge themselves on process measures — follow-up within seven days, a visit after discharge — that are cheap to compute and have little demonstrated relation to recovery. Using linked person-level claims and outcome records, the paper asks how much of the variation in whether anyone actually got better these measures explain, and what a measure built on observed outcomes would look like instead.