All writing

Bringing Decision Models into Knowledge Work

Scaling System 1 with low-cost professional judgment

Knowledge work produces reports, legal opinions and investment conclusions, yet most of the work happens before anything is written: deciding which materials are relevant, which clause applies, which number contradicts what management said. These judgements decide where attention goes and what the conclusion rests on.

Professional judgement has two properties. First, many judgements rest on recognising a situation and on professional intuition rather than on lengthy reasoning; an experienced lawyer reading a clause usually knows quickly whether it is a restriction. Second, judgements vastly outnumber conclusions; behind a single investment conclusion lies the screening of many documents, entities and points in time.

The main route to more capable models in recent years has been larger parameter counts and longer reasoning, that is, scaling System 2. It raises the ceiling of each judgement, and also its cost, so coverage is bounded directly by budget. This post discusses another route: provided each judgement is reliable enough, cut its cost by one to two orders of magnitude and thereby expand the number and coverage of judgements. We call this scaling System 1.

We use Jev as a case. Jev is a model that outputs judgements as choices, scores and probabilities. 1 Using public legal and financial data, we take up three questions in turn: how far judgement can go without long reasoning; how much the same budget covers once judgement is cheap; and where long reasoning should be spent. Two cases then show how these judgements enter real research.

1. Professional Judgement Does Not Require Long Reasoning

We ran Jev and two settings of DeepSeek V4.1 Flash on the same items, one call per item. Below, DeepSeek is abbreviated as DS and its two settings are labelled “no thinking” and “medium”.

BenchmarkJev
native
DS
no thinking
DS
medium
MAUD77/9070/9079/90
ContractNLI121/170125/170125/170
FinDVer-KNOW49/6038/6046/60
FinArgQuality46/9650/9651/96
Consumer Contracts QA32/3231/3232/32
Median latency per item0.4–0.5 s1.1–1.3 s1.8–9.5 s
Total cost, five tasks$0.044$0.911$1.904

Tasks: MAUD covers judgements on 9 legal standards in merger agreements; ContractNLI is inference over NDAs; FinDVer-KNOW is knowledge-intensive verification of claims in financial reports; FinArgQuality rates earnings-call argument quality, 24 arguments × 4 dimensions, with a per-dimension calibration fixed in advance for Jev; Consumer Contracts QA is the LegalBench task on consumer-contract rights and obligations (task name consumer_contracts_qa), 32 distinct excerpts with one question each. Scores are correct / judgements, not full dataset sizes.

Settings: “no thinking” is reasoning_effort=none and “medium” is reasoning_effort=medium; Jev uses its native structured-judgement interface. The accent colour marks the Jev column only, not the best result in each row. Latency is the range of per-task median latencies, not batch time.

Judgement ability need not come attached to a large model. With neither side using long reasoning, Jev and DS (no thinking) each win some tasks and lose others, and are at the same overall level; Jev costs about 1/20 as much. This suggests that, at least on these tasks, judgement quality depends more on whether a model has been shaped into a judge, giving comparable scores over constrained options, than on how much general knowledge its parameters store.

The return on long reasoning depends on the structure of the task. With thinking on, DS gains 9 items on MAUD and 8 on FinDVer, but almost nothing on ContractNLI and FinArgQuality. The first two require checking clauses against legal standards and reconciling financial definitions step by step, so the reasoning itself helps; the latter two lean on semantic intuition and subjective weighing, where thinking longer is not necessarily more accurate. Even on MAUD and FinDVer, where reasoning helps most, Jev reaches similar scores at 1/46 and 1/37 of the cost of DS (medium). On FinArgQuality, Macro-F1 is .376, .352 and .423 respectively.

We also built a set of professional trade-off tests that require choosing among constraints, timing and cost. Each model is shown in the fast configuration actually used:

ModelReasoning settingPrimary action match
Jev 1.13.0Native judgement5/9
DeepSeek V4.1 FlashNo thinking6/9
GLM-5.3-FlashLow6/9
Gemini 3.6 FlashNo thinking7/9
Hunyuan 3No thinking5/9
Cogito-32BDirect answer5/9
Qwen3-30B-A3BDirect answer1/9
BERT-MNLIEntailment2/9

Task: 9 author-constructed scenarios in finance, law and medicine, each with 12 records and 5 candidate actions.

Settings: DS uses none, GLM low, Gemini thinking budget=0, Hunyuan no_think. BERT turns each candidate action into a textual-entailment judgement.

The fast settings differ little from one another; the clear break is the BERT row. It is just as fast, but can only judge whether one text entails another and cannot weigh several constraints. The value of fast judgement lies in keeping the ability to make trade-offs while being fast, which is exactly what separates professional judgement from text matching.

2. Lower Cost, Wider Coverage

We gave the 60 FinDVer items to each model in a fixed order and accumulated cost item by item, tracking how correct judgements grow with spending.

FinDVer-KNOW: correct judgements versus cumulative cost

Figure 1 · FinDVer-KNOW, 60 items, fixed order. The x-axis is cumulative estimated cost (log scale); the y-axis is cumulative correct judgements.

The three curves have similar shapes, so per-item accuracy differs only modestly; what separates them is the x-axis. At $0.01, Jev has processed 47 items and answered 38 correctly, while DS (medium) has processed 1.

Under a budget, the more meaningful question is how many correct judgements a given spend buys; per-item accuracy is only one factor. The objects of research have no natural upper limit: documents, entities and points in time can keep growing. When the cost of a judgement falls by an order of magnitude, material that was only skimmed in summary can be read in full, and objects that were sampled can be checked exhaustively. This is the direction of scaling System 1.

3. Concentrating Long Reasoning Where Confidence Is Low

Scaling System 1 is meant to let System 2 concentrate where it matters. Jev outputs a confidence for each judgement. 2 Under a fixed quota, we sent the lowest-confidence items to DS, with the escalation list fixed before any DS call.

Benchmark
escalation
Kept by Jev
items · accuracy
Sent to DS
items · before → after
Jev + routing
correct
All DS
correct
Est. cost
routed / all
MAUD
medium
67
92.5%
23
65.2% → 78.3%
80/9079/90$0.105 / $0.306
ContractNLI
no thinking
130
83.8%
40
57.5% → 70.0%
137/170125/170$0.147 / $0.458
FinDVer-KNOW
medium
45
86.7%
15
66.7% → 66.7%
49/6046/60$0.163 / $0.467

Reading the table: tasks are defined under Table 1. “Kept by Jev” is the part not escalated; “before → after” is Jev → DS accuracy on the same escalated items. “All DS” uses the setting given in that row; costs are in US dollars and include all first-pass judgements and any follow-up calls.

Confidence concentrates errors in a region that can be identified. On all three tasks, the items Jev kept are clearly more accurate than those it passed on, so long reasoning only needs to cover about a quarter of the items. The two models also err in different places: on MAUD, there are 5 items Jev gets right and DS (medium) gets wrong, and 7 the other way round. Routing keeps the strengths of both, and on all three tasks the final score is no lower than either model alone. The ContractNLI cascade was a full live run that took 25.1 s for the whole batch; a separate run of all-DS (no thinking) took 71.9 s.

In this division of labour, System 1 decides where it is worth thinking harder, and System 2 thinks those places through. The cheaper the judgement and the more reliable its confidence, the more expensive reasoning can be concentrated on the problems that are actually hard.

4. Scaling Judgement in Research: Two Cases

The two cases below show two uses of low-cost judgement in research: organising evidence within a single question, and tracking hypotheses over time.

3M: EPS rose, but did sustainable profitability improve?

In 3M’s 2022 annual report, GAAP EPS rose from $10.12 to $10.18, while adjusted EPS fell from $10.73 to $10.10 and operating cash flow declined. The research question: did the rise in EPS come from an improvement in ongoing operations?

We extracted 17 passages from the annual report, covering profit, adjusted measures, cash flow, litigation and environmental obligations, and paired them all, giving 136 Jev relationship judgements: should these two passages be read together, and does one qualify the other? DS (medium) then went back to the source text and wrote the report from this list of relationships. Because each judgement is cheap enough, every pair can be compared, with no need to guess in advance which materials are related.

EPS and cash flow diverge → Jev organises evidence relationships → the report asks: how much of the earnings decline comes from raw materials and currency, and how much from the business itself?

The report thus turns “earnings quality is weakening” into a testable question: if the decline comes mainly from temporary factors, the conclusion can be softened; otherwise the remaining businesses deserve a closer look. Jev’s output is a list of questions worth checking; the conclusion still comes from the synthesis model and the researcher.

CVS: can a new disclosure bring an old judgement back into the research queue?

This case is about long-term tracking. The researcher maintains a set of working hypotheses; whenever a new disclosure arrives, Jev judges whether it bears on a hypothesis and whether the researcher should be prompted to review it.

At the end of 2020, CVS still had about $13.9 billion of unused buyback authorisation, yet made no repurchases in the fourth quarter, so the research kept the hypothesis “is the company prioritising capital preservation?”. The 2022 annual report then disclosed about $3.5 billion of actual buybacks.

Old hypothesis: preserving capital? → New evidence: about $3.5 billion actually repurchased → Jev alert: buybacks have resumed; the capital-allocation hypothesis should be updated.

This alert requires distinguishing “authorisation” from “execution” and then connecting the new disclosure back to the old hypothesis. The workload of long-term tracking is the number of new materials times the number of hypotheses being tracked, and it keeps accumulating over time. That product soon exceeds any budget for long reasoning, yet suits low-cost judgement well. In a tracking experiment of eight cases with 96 dated updates each, all 12 alerts Jev issued were supported by their sources.

5. Outlook: Scale, Capability and Access

Scaling up judgement. The experiments here are measured in tens to hundreds of items. Real professional work faces entire document collections, continuously updated disclosures and spans of many years. When judgement reaches that scale, the question shifts from “is each judgement correct?” to “did we, overall, see what we should have seen?”.

Strengthening professional judgement through targeted training. The experiments suggest that professional judgement does not necessarily depend on parameter scale and long reasoning, and can be built as a capability in its own right. Training on specific types of judgement, such as distinguishing legal standards, reconciling financial definitions and assessing argument quality, with expert annotation and structured judgement tasks, could keep raising judgement quality without raising the cost per judgement, so that more problems can be handled reliably without escalating to long reasoning.

Making professional judgement widely accessible. Today, sustained and careful professional scrutiny of large volumes of documents is mostly affordable only for large institutions. When reliable judgement is cheap enough, individual investors, small law firms and independent researchers can also read more widely, follow longer and spot problems earlier. The division of labour we hope for: fast judgement spreads wide, long reasoning digs deep, and people make the final decisions.


References

  1. TypeSafe. Jev: Models and Pricing.
  2. TypeSafe. Confidence versus Probability.
  3. MAUD: Merger Agreement Understanding Dataset.
  4. ContractNLI: Natural Language Inference for Contracts.
  5. FinDVer: Claim Verification over Financial Documents.
  6. FinArgQuality: Quality of Managers’ Arguments in Earnings Conference Calls.