Read the amounts, but left the numeric fields blank.
Preserved £9.18.9 and £3.15.0 in a text note. It also recognized that these duties stayed unchanged but did not put either amount into the numeric duty columns.

LLMS + UNCONVENTIONAL DATA
These datasets will be digitized once and used for academic research. With limited human review, quality is a top priority, and checks must be built into the workflow. I follow three principles:
For a project digitizing international trade negotiation records with human transcriptions available for a subset of the data, my panel workflow reduced clear errors from 13.4% to 1.5%, an 89% reduction, compared with an earlier Claude-only workflow.
See below some LLM projects in which I have had a central role. You can learn more about the data problem, the solution, some cool illustrative examples, the added value of my approach, and the general workflow structure.
My role. I lead construction of the dataset. I designed the LLM panels for digitization and product classification and evaluate their outputs against reference records.
The project draws on 100,000+ pages of declassified trade negotiations spanning 50+ years. Complex, changing page layouts that are unknown in advance present requests, offers, concessions, and different tariff types for different products. Each country used its own product classification. To compare bargaining across countries and over time, I first need to recover what each record says, structure the information in a dataset, and link the products to a product classification that is common across countries.
I designed two LLM workflows. The first recovers the rows on each page, then asks three models to read small groups of products independently. Their proposed fields, confidence scores, and reasoning go to a final model that rereads the page and decides what belongs in each of the final structured dataset’s 39 columns that measure dimensions of interest (for example, tariff type: fixed, ad valorem, etc.).
The second assigns modern, standardized six-digit product codes (HS6) to each product description. GPT and Gemini review the same product independently. Agreed codes are kept; Claude reviews disagreements, and a higher-tier GPT model handles cases that remain unresolved. I ask the different models to output their reasoning and later use it to help later-stage reviewers make a confident decision. I also created a rubric to have them score their confidence in their decisions: cases that still lack a conclusion or fall below the rubric’s confidence threshold go to human review.
WHAT THIS MAKES POSSIBLEA product-level record of what governments requested, offered, and agreed to, linked to modern trade classifications. This makes it possible to study how governments bargain over protection across countries and over time.
Australia offered Finland a cut in a surcharge on paper products. Recovering the offer meant reading old currency, separating two tariff schedules, and working out which duties stayed in force.

The page has separate standard and preferential duties, each with a per-ton charge and a surcharge called “primage.” “Exempt” refers to that surcharge, not the entire tariff. The amounts use pounds, shillings, and pence.
Open the scan at full size ↗01 THREE INDEPENDENT READINGS OF THE SAME ROW
Note that Australia used the pre-decimal pounds, shillings, and pence system at the time, so the models exploited their knowledge of the historical context to convert these amounts to decimals.
Preserved £9.18.9 and £3.15.0 in a text note. It also recognized that these duties stayed unchanged but did not put either amount into the numeric duty columns.
Correctly reasoned that cutting the surcharge leaves the per-ton charges in place. But it read the preferential duty as £9.15.0 instead of £3.15.0, recording £9.75 where the value should be £3.75.
Correctly converted the existing duties to £9.9375 and £3.75 per ton. Its explanation said they remained unchanged, yet the proposed-duty columns were left blank.
02 REREAD THE PAGE WITH ALL THREE REVIEWS
The final review checked the scan, confirmed £3.15.0, and converted both amounts exactly. It used the request and response columns to establish that only the surcharge changes, then filled the unchanged per-ton duties into the proposed-rate columns.
| Duty component | Before | Australia’s offer |
|---|---|---|
| Standard · per ton | £9.9375 | £9.9375 unchanged |
| Standard · surcharge | 10% | 5% |
| Preferential · per ton | £3.75 | £3.75 unchanged |
| Preferential · surcharge | 5% | 0% exempt |
The unchanged duties are recorded as derived from the offer’s wording. The original amounts and the reason for carrying them forward remain attached to the record.
Source: Australia–Finland tariff negotiations, 1949, file 500083-0004.pdf, physical page 4. The standard and preferential columns are headed M.F.N. and B.P.T. The old currency converts as follows: £9.18.9 = 9 + 18/20 + 9/240 = £9.9375; £3.15.0 = 3 + 15/20 = £3.75.
The archived identity ledger identifies the independent readers as Gemini 3.1 Pro, GPT 5.6 Sol xhigh, and Claude for this case. The final page review used GPT 5.6 Sol with Ultra reasoning. The explanations above summarize their saved outputs; they are not invented quotations. The final values were also checked against the exported dataset.
This example demonstrates a documented digitization improvement. The separate 13.4% to 1.5% comparison above comes from an audit against human transcriptions, not from this page.
Read the saved decisions and final record ↗A blood-cell counter. Three different proposed mappings. A final review that identifies the instrument and preserves its parts.
Optical instruments: Haemacytometers, frames, mountings, and parts
A haemacytometer counts cells in a sample. The task is to translate this historical description into comparable six-digit product codes. Does “optical” determine the category, or does the instrument’s function? And where do its parts belong?
See the original page ↗01 TWO INDEPENDENT REVIEWS
Considered medical, analytical, and optical instruments, alongside general parts categories. Its six-code answer did not include the specific laboratory-parts code.
901890902780903140903180903190903300Included the laboratory instrument and its parts, but also medical instruments, optical measuring equipment, and apparatus using optical radiation.
90189090275090278090279090314090319002 REVIEW THE DISAGREEMENT
Claude narrowed the proposals to five codes. It recognized that these were competing interpretations of one instrument, not five different products. The key question remained: medical equipment, laboratory analysis, or optical measurement?
90189090278090279090314090319003 FINAL REVIEW, USING THE EARLIER DECISIONS
The final review treated the haemacytometer as an instrument for analyzing laboratory samples. That supports the analytical-instrument category over general optical or medical equipment. Because the record explicitly includes frames, mountings, and parts, the corresponding parts code belongs in the mapping too.
Source: United States–Germany negotiations, scanned page 69 (printed page 10), tariff item 228(a), statistical code 9180.020. The display uses the saved GPT and Gemini reviews after candidate expansion, followed by Claude’s arbitration and the final GPT review. Explanations above are plain-language summaries, not invented quotations.
The saved model identifiers are GPT 5.5, Gemini 3.1 Pro (High), Claude Opus 4.8, and GPT 5.5 for final escalation. The final decision retained 902780 and 902790 under the workflow’s HS 1988/1992 reference. It records a confidence score of 70/100 and no human-review flag; this score is a model assessment, not measured accuracy.
This is a documented classification decision. The 13.4% to 1.5% error comparison above concerns a separate tariff-extraction audit against human transcription, not this HS6 mapping.
Read the saved model explanations and code sets ↗The common comparison contains 59,435 populated tariff fields. The earlier workflow recovered 44,714, left 6,758 unresolved or otherwise uncredited, and made 7,963 clear errors. The panel recovered 54,530, left 4,029 uncredited, and made 876 clear errors. The reduction in clear errors is 89%, or roughly ninefold.
Exact values and source-supported equivalents receive credit. Clear omissions, wrong values, and values placed in the wrong tariff field count as errors. Ambiguous readings, apparent reference problems, and other uncredited cases remain separate. The difficult-case audit is model assisted and uses retained source evidence; it is not a fresh human transcription of every field. The results apply to this comparison subset, not to every project or to product classification.
Model assignments vary across runs. This diagram shows the general workflow; the worked examples identify the models used in each saved record.
Page images → clean tariff rows
Old product wording → comparable HS6 codes
Shared rubric: reviewers choose only from the supplied HS6 candidates, support the decision with the historical wording and context, preserve one-to-many mappings when needed, and report confidence on a 0–100 scale.
My role. I conceived and led this study of Korea’s R&D push, including data construction, empirical and structural analyses, and writing.
South Korea’s G7 Program, an ambitious public R&D program, left 4,771 detailed descriptions of government-sponsored research projects in Korean. They describe each project’s goals and activities but do not identify the corresponding patent or product classes. Without those links, I cannot compare targeted technologies with the patent and export data needed to evaluate the policy.
I run the same review workflow twice for every project: once to assign patent classes and once to assign traded-product classes. GPT 5.6 Sol and Gemini read the description, goals, and activities independently. Each records its classification, reasoning, and confidence using a shared rubric. Agreements are kept; disagreements go to DeepSeek V4 Pro with the earlier reasoning. Unresolved cases receive a final read from GPT 5.6 Sol Ultra. Anything still unclear or below the confidence threshold goes to human review.
WHAT THIS MAKES POSSIBLEThe classifications turn the original project descriptions into a treatment measure that can be linked to patents and exports. Compared with selected-but-never-funded technologies, technologies targeted by an ambitious government R&D program more than doubled citation-weighted patenting and nearly tripled exports within ten years, with benefits exceeding costs threefold. Read the paper’s findings ↗
The title names gene therapy. The objective specifies cancer treatment and clinical use. The activities describe production, purification, and a production system. The task is to identify which technologies the project develops, not simply assign a code to every technical term.
English translationDevelopment of a production process for adenoviral vectors for gene therapy.
Original Korean유전자 치료 용 아데노 바이 러스 벡터 의 생산 공정 개 발
English translationDevelop mass-production technology for adenoviral vectors suitable for clinical cancer treatment, and apply them clinically.
Original Korean임상 적용 가능한 암 치료 용 아데노 바이 러스 벡터 의 대량 생산 기술 개발 및 임상 적용
English translationOptimize pilot-scale production and purification methods for adenoviral vectors. Establish a mass-production and purification system. Optimize adenoviral-vector production in a GLP facility.
Original Koreano 아데노 바이러스 벡터 의 pilot - scale 생산 방 법 및 정제 방법 의 최적화 대량 생산 및 정제 system 구축 OGLP facility 하에서 아데노 바이러스 벡터 의 생산 최적화
The title, objective, and activities are shown in full. The models read the Korean record; the English translation is provided for readers.
01 TWO INDEPENDENT REVIEWS
Selected codes for viral-vector production, genetic engineering, gene therapy, and cancer treatment. It also interpreted “establish a mass-production and purification system” as support for virus-culture equipment, although no equipment design was specified.
C12N 7/00C12N 15/00A61K 48/00A61P 35/00C12M 3/00Treated cancer gene therapy as the intended use rather than a separate technological output. It selected the two viral-vector codes and proposed two equipment codes, with low confidence because the activities emphasize methods rather than apparatus.
C12N 7/00C12N 15/00C12M 3/00C12M 1/0002 REVIEW THE DISAGREEMENT
Read the clinical objective as evidence for gene therapy and cancer treatment, alongside the production process. It kept the virus-culture equipment code but acknowledged that the record described no apparatus structure or design. That low confidence led to a final review.
C12N 7/00C12N 15/00A61K 48/00A61P 35/00C12M 3/0003 FINAL REVIEW, USING ALL EARLIER POSITIONS
Clinical application is part of the stated objective, so the treatment is more than a possible downstream use. But establishing a production system does not, by itself, describe a new reactor, culture vessel, or purification device. The final review kept four codes and excluded the equipment classifications.
Project 1999_3981, classified in July 2026. The complete title, final objective, and technical activities appear above. The saved model input also contains a repeated technical-text block, preserved in the evidence file. The English translation normalizes spacing and list markers; “OGLP” is read as a bullet followed by “GLP,” as in the saved reviews.
The actual reviewers were GPT 5.6 Sol XHigh and Claude Opus 4.8 XHigh independently, Gemini 3.1 Pro High for arbitration, and GPT 5.6 Sol Max for the final review. These are the model assignments recorded for this case, which can differ across runs.
Neither initial reviewer nor the arbitrator returned the final four-code set. All three proposed C12M 3/00, the virus-culture apparatus group; the final review excluded it. The final exported classification matches the four codes displayed above, using the frozen IPC 2006.01 scheme.
The explanations summarize saved decisions, not newly generated responses. The final review recorded a rubric confidence of 0.91 and no human-review flag. This is a documented classification decision, not an independent accuracy test.
Read the full Korean source, English translation, and saved decisions ↗This configuration uses GPT and Gemini, then DeepSeek and a final GPT review. The saved gene-therapy example above used GPT and Claude, then Gemini and a final GPT review. The review structure is shared; the model assignments differ by run.
THE ORIGINAL TEXT, EACH MODEL'S ANSWER, AND ANY UNCERTAINTY ARE SAVED WITH EVERY LINK
My role. I am building a new technology taxonomy from patents and academic papers and developing Bayesian models to forecast which technologies will emerge.
Before forecasting a technology, I need to identify it consistently over time. Existing patent and publication classifications often group distinct technologies together, obscuring new ones. I need a more detailed taxonomy built from patents and papers, using only the information available at each historical date.
I am building the taxonomy with LLMs that identify technical concepts and organize the supporting papers and patents. I use this evidence to measure dimensions of emergence identified by the science-of-science literature, such as growth and knowledge flows across fields. I am developing Bayesian models to use these data to forecast which technologies will emerge over specified time horizons.
WHAT THIS MAKES POSSIBLEThe goal is to identify emerging technologies before their importance becomes obvious. Rebuilding the data at past dates will let me compare forecasts with what happened later, without giving the forecast information from the future.
LLMs recover precise technical concepts from each source. Consistent rules link concepts, documents, and citations into technologies that can be reconstructed over time.
For each technology and year, construct the dimensions that the science of science literature identifies as signals of emergence.
A hierarchical Bayesian model will combine these measures and learn from comparable technologies.