All writing

Scaling System 1 towards Inclusive Intelligence

Introducing Jev Cowork

Much of the cost of knowledge work arises before an answer appears: identifying the facts that matter, allocating attention among competing explanations, and deciding what to investigate next. Even after a report is finished, a new disclosure, court ruling, or experimental result may invalidate one of its judgments. The research must continue.

This work depends on a kind of professional intuition: knowing what deserves attention, what warrants doubt, and what is sufficient. We call it taste. Jev delivers judgments directly as selections, scores, and probabilities. Its low cost and low latency create an opportunity to bring this intuition into more parts of the research process.

Jev Cowork operates along two dimensions: within a question, broadening the coverage of evidence and judgments about relationships; over time, maintaining assumptions, incorporating changes, and advancing investigations. We call this direction Inclusive Intelligence: allowing more information to be considered and more questions to receive affordable, sustained research attention.

1. Judgment and prediction: two capabilities of System 1

In this article, System 1 refers to model calls that directly produce bounded judgments without an explicit chain of thought. We tested the structured interface of jev-1.13.0.[1]

System 1 describes how a capability is invoked: quickly and directly, without constructing a full argument every time. It can be used for either prediction or judgment. The distinction between these capabilities comes from the task itself.

Prediction concerns outcomes: given the information available, what will happen, and with what probabilities? Judgment concerns choices: given goals, evidence, and constraints, what deserves belief, attention, or action? We test the latter through action selection, then extend it to the allocation of attention in research.

The two draw on shared knowledge but have different success criteria. Someone may correctly anticipate growing demand yet choose the wrong financing plan. Conversely, someone unable to predict market movements precisely may still recognize that a proposal violates a liquidity constraint, or identify the evidence that should be checked first. Understanding the world and choosing how to act require separate evaluation.

CapabilityCore questionPrimary evaluation criteria
PredictionWhich outcomes will occur, and with what probabilities?Probability error and calibration after outcomes are known
Judgment, tested here through action selectionWhat should be chosen under these goals and constraints?Constraint satisfaction, trade-offs among objectives, and action quality

Output length does not distinguish them either. A single probability can encode a complex prediction; a single option can encode a complex professional judgment. Taste is especially apparent when several choices all seem reasonable: which conditions cannot be violated, which exception actually applies, and whether a feasible plan deserves priority.

Judging actions in a given situation

Hard Pro contains nine cases across finance, law, and medicine. Each question includes twelve records and five candidate actions, requiring the model to integrate constraints across records, temporal relationships, and the costs of different plans. The questions and author-defined reference criteria were frozen before the runs. Models used neither retrieval nor tools. The results are shown below.

Hard Pro: Jev, Hunyuan 3, and Cogito direct each score 5/9; Gemini 3.6 Flash scores 7/9; Qwen3 raw direct scores 1/9; BERT-MNLI scores 2/9.
Figure 1. Primary-action matches on the same Hard Pro question set.

Jev matches Hunyuan 3 no_think and Cogito-32B direct, scoring one question below DS none and GLM low. Gemini 3.6 Flash scores 7/9. BERT-MNLI converts each action into a separate entailment judgment and scores 2/9. There is a clear gap between local semantic matching and professional action selection: the latter requires organizing conditions scattered across the material into a coherent whole.

Harder finance questions also expose the limits of this intuition. Jev and DS none both score 0/3, while DS max scores 3/3. In one financing question, Jev chooses a feasible but more expensive plan. Recognizing local conditions and resolving the overall trade-off are different levels of capability.

Latency reveals another difference. The median duration of Jev's primary-action requests is 0.846 seconds, compared with 1.887 seconds for DS none and 9.004 seconds for DS max. These are observations from the experimental pipeline. For local judgments that must recur frequently, subsecond responses already have practical value.

Probabilistic prediction of unknown outcomes

Beyond action selection, we also tested historical-event forecasting. On the seven financial questions for which all six configurations completed successfully, the Brier scores are as follows. Lower is better.

ConfigurationBrier score
DS V4.1 Flash max0.05347
Gemini 3.6 Flash, thinking budget 00.09671
GLM-5.3-Flash low0.11794
DS V4.1 Flash none0.13233
Hunyuan 3 no_think0.14357
Jev0.15239

Jev trails the other configurations on this forecasting set. Together with the action-selection experiment, this gives a more specific capability profile: it can handle some professional constraints and action trade-offs, while its probability estimates for unknown events are comparatively weaker. The two experiments measure different aspects and cannot substitute for each other.

The same distinction applies to System 1. Quickly recognizing a situation's key structure and quickly producing calibrated probabilities require learning different mappings. Directly choosing a suitable action does not guarantee accurate probabilities over future outcomes. Even when an interface returns confidence, its calibration still needs to be evaluated separately.[2]

In another twelve-question experiment, having DS first generate a summary, a mechanistic interpretation, or an evidence audit and then passing it to Jev did not reduce mean Brier error relative to using Jev directly. What an explanation adds and whether that information improves probability estimation are separate questions.

Combining judgment and prediction

Real decisions often require both prediction and judgment. Prediction provides possible outcomes; judgment determines how to weigh gains, losses, and constraints. The same outcome probabilities can lead to different actions under different funding horizons or costs of failure. If an action changes the outcome, the prediction itself must also be conditioned on that action.

Hard Pro places rules and key conditions in the question, allowing us to observe trade-off decisions relatively directly. Open-ended research requires an earlier judgment: what information is missing, which uncertainty matters, and how much investigation is worth spending on it. Here, taste helps determine where prediction and reasoning should be applied.

This suggests a concrete role for Jev within a system. Even if it cannot yet accurately forecast a company's profit next year, it may still identify an important disclosure, notice tension between two accounts, or flag an assumption worth revisiting. The following experiments test whether these judgments can improve the research process.

2. The marginal cost of judgment

When a professional judgment becomes cheap enough, new ways of organizing research become possible.

At the experimental list price of $0.042 per million input tokens, ten thousand judgments with two thousand tokens each cost approximately $0.84 in input charges. One set of Jev Cowork research experiments made 934 independent Jev calls, using approximately 2.56 million input tokens at an estimated input cost of $0.108. Synthesis-model, retrieval, and verification costs are additional.

At this scale, we can make the objects of judgment more granular: the importance of a source, the complementarity of two pieces of evidence, the scope of a conclusion, or the effect of a new disclosure on an old assumption. These judgments can be computed separately, run in parallel, and retained as research state.

What scaling expands here is coverage across distinct judgments. Asking the same question ten times may leave the same blind spot intact; identifying ten new evidence relationships may open another path of investigation. The value of scale depends on the information gained from the additional judgments.

Knowledge work is always constrained by attention. Researchers prioritize a few urgent questions, while other uncertainties, weak signals, and long-term assumptions must wait. Low-cost judgment offers a way to broaden this attention: more material can be noticed, more connections can be proposed, and questions that nobody is actively rereading can continue to receive new information.

Inclusive Intelligence has two meanings. Within research, more evidence and explanations get a chance to be considered. Beyond an individual research task, lower maintenance costs may allow small teams and individuals to sustain attention across larger sets of questions. Jev Cowork explores these possibilities through deep research and continuous research, respectively.

3. Going deeper: broadening evidence and relationship coverage

Jev judges which materials deserve to be read together and where qualifications or counterevidence arise. Code organizes these relationships into an agenda, and a synthesis model combines it with the full source text to produce a cited report.

JEV COWORK / DEEP RESEARCHFrom evidence to a research brief
Source materials & research question
Jev relationship judgmentsImportance · Complementarity · Scope
Research agendaDeduplication · Coverage · Composition
Full source contextOriginal text, table headers and scope
Synthesis model → Cited research briefResearch agenda + full source context
Source verification & human judgment

From finding material to building an argument

The evidence-retrieval experiment covers 380 complete documents and 90,041 source passages.

Four retrieval metrics: financial target-page localization, financial document ranking, legal document ranking, and legal positive-evidence coverage.
Figure 2. The scorable finance set contains 17 questions across 4 companies, weighted equally by company; the legal set contains 13 cited cases, weighted equally by case. Hit@20 measures whether the target page appears among the top twenty results, MRR measures the rank of the first relevant document, and Recall@20 measures positive-evidence coverage.

Financial target-page Hit@20 rises from 0.107 with lexical retrieval to 0.929 with Jev ranking; DS max scores 0.893 in the same round. Jev also leads on legal document ranking, but DS max still leads on positive-evidence coverage.

Finding material is not the same as completing an argument. Across six research questions on SVB and Akorn, Jev evidence composition plus DS max receives a source-verification score of 78.33/100, versus 90.83/100 for DS max given the complete evidence directly. Filtering can discard definitions and table headers. Report generation therefore preserves the full source text, letting the agenda guide reading rather than replace the sources.

Intuition proposes connections; sources constrain conclusions

For four research tasks on 3M and Waterkeeper, 512 distinct Jev relationship judgments form an agenda, which DS medium uses alongside complete source pages to write the reports.

3M / Waterkeeper, four tasksJev agenda + DS mediumDS max aloneDS none alone
Research-task success3/43/42/4
Key facts verified20/2020/2015/20
Predeclared major errors001
Audited citations supported by sources60/6157/5938/49
Writing output tokens, including reasoning49,99367,51211,137
Mean writing time55.9 seconds72.4 seconds14.6 seconds

Citations underwent model-assisted source verification. Token counts are totals across the four tasks; time is the mean duration of a single DS writing request. Agenda construction and research evaluation are accounted for separately.

The agenda-plus-medium configuration matches max alone on task success and key-fact verification. Each configuration has one legal report that fails delivery checks because of citation-number errors. The former uses 25.9% fewer writing output tokens and takes 22.7% less time on average. Building the four agendas with Jev costs an estimated $0.0624 at list price. These results reflect the combination of an agenda and a lower reasoning setting.

Beyond these numbers, the reports also leave questions for further investigation. At 3M, earnings per share rise slightly while adjusted earnings and cash flow weaken. Rather than stopping at "declining earnings quality," the report considers raw-material and exchange-rate effects alongside volume, pricing, and productivity. It asks how much of the decline might ease with external conditions, and how much stems from the business itself. If temporary cost pressure is the main cause, the assessment of sustained earnings may become less negative. If volume and unit profitability also continue to weaken, the operating performance of the remaining businesses needs closer examination. Follow-up research thus has concrete targets, rather than merely an instruction to "keep watching the financial reports."

In Waterkeeper, the connection comes from different pages of the same judgment. One passage says that a party did not concede the waters' relative permanence but did not directly dispute that fact either; another says that it challenged the opposing party's assertion of "permanent standing waters." Reading these passages together, the report proposes a line of inquiry: return to the pretrial statements and trial record to establish which physical facts were conceded and which remain disputed. This bears on what additional evidence is needed under the new legal standard and which questions subsequent proceedings can address.

The information gain comes from connecting existing evidence. No new source material is added, but scattered facts are organized into explanations that need to be distinguished and premises that need to be established. Here, the agenda provides research leads, not answers. The next step still requires ordering the work: which question is most likely to change the conclusion, which checks depend on earlier results, and what evidence would justify stopping.

4. Over time: keeping research active

Judgments in reports are often conditional: "the company is still prioritizing capital preservation," "the conclusion depends on a regulatory condition," or "next quarter's cash flow needs watching." These conditions are rarely maintained separately. When new information arrives, a researcher must recover the old judgment and reconnect it to the evidence.

Jev Cowork saves findings from briefs as working assumptions, records conditions for review, and continuously matches subsequent material against them. Changes worth attention generate alerts; evidence obtained through investigation enters the same research archive.

JEV COWORK / CONTINUOUS RESEARCHKeeping judgments connected to evidence
Research findings
Working assumptionsJudgments and review conditions
New source materialsNew disclosures and evidence
Jev: does this merit review?Connect new evidence to active assumptions
Source-linked review alerts
Human review & further investigation
Updated evidence, assumptions & brief
↳ Continue maintaining working assumptions

CVS: updating a capital-allocation assumption

At the end of 2020, CVS still had substantial unused share-repurchase authorization but made no repurchases in the fourth quarter. A researcher could retain a capital-allocation assumption: is the company still prioritizing capital preservation?

The subsequent 2022 annual report discloses approximately $3.5 billion in actual repurchases. The system connects this new material to the earlier assumption and proposes a review alert: repurchases have resumed, so the capital-allocation judgment deserves updating.

English rendering of the saved Jev Cowork CVS case: a working assumption on the left and a review alert triggered by a new disclosure on the right.
Figure 3. Translated and reformatted from the saved CVS case: from a capital-allocation assumption to an alert about resumed repurchases. Both Jev and DS identify this update, and the alerts receive source support in model-assisted review.

The basic unit of research thus expands to relationships among questions, assumptions, and evidence. A judgment left by one report can continue to participate in subsequent research.

Continuous monitoring and investigative actions

The continuous-research experiment covers eight corporate or legal topics, backed by 397 complete documents. Each case contains 256 candidate materials and 96 chronological updates. Attention allocation, continuous monitoring, and investigative actions are tested separately.

Continuous monitoring: Jev source-supported alerts 12/12 and mean timely recall 37.5%; DS 7/12 and 41.7%; rules 2/18 and 4.2%.
Figure 4. Left: model-assessed source support for delivered alerts. Right: mean timely recall across eight cases. Jev covers 5/20 review targets in time; pooled recall is 25%, while the equally weighted case mean is 37.5%.

All 12 alerts delivered by Jev receive source support, compared with 7/12 for DS and 2/18 for rules. The higher alert quality makes this approach worth exploring. However, Jev's mean timely recall across cases is only 37.5%, leaving many omissions. At present, it is suited to assisting researchers with ongoing attention.

Attention allocation and investigation depend more heavily on the specific workflow design. Among the six attention-allocation cases that passed the control checks, Jev fully supports 13/24 questions, versus 12/24 for rules. In the investigation task, Jev scores 18/32, while rules reach 20/32. The Block case contains a path to complete evidence collection within budget, but Jev's actual actions sufficiently support only one of the four questions.

This exposes a deeper requirement for continuous research: the system must remember what has been resolved, what is still missing, and choose the next step accordingly. Judgments need shared state. Otherwise, even when every call is cheap, many local choices may repeatedly expend effort on the same question.

5. Jev Cowork: a collaborative research workspace

Jev Cowork is currently a research prototype centered on case demonstrations. It organizes the methods above into an interactive workspace for replaying existing results, observing research workflows, and trying new material. The current implementation is built primarily around the cases in this article; generality, stability, and the experience of long-term use still need improvement.

Project repository: Yii-Jing/Jev-Cowork.

The repository connects deep and continuous research through a unified research archive, linking material import, evidence organization, and assumption maintenance to demonstrate the basic form of this collaboration.

The workspace contains five connected views:

ViewPrimary function
Deep researchInspect evidence relationships and organize a research agenda
Evidence queueRetrieve source text and track reading and verification status
Research briefRead cited reports alongside their sources
Continuous researchMaintain working assumptions and receive updates and review alerts
Investigation and notesSave follow-up leads and research notes, and export archives

The repository includes two complete examples. The 3M example contains 136 sets of saved Jev relationship judgments, allowing comparisons between agenda-guided reports and reports generated through direct synthesis. The CVS example shows how capital-allocation assumptions connect to subsequent disclosures. Both examples can be replayed locally to inspect the correspondence among source text, judgments, and reports.

Users can import TXT, Markdown, JSON, or PDF material to create their own research archives. Materials, notes, and run records are stored locally. Archives support Markdown and JSON export, as well as restoration from JSON. Live model analysis requires the user's own Jev API key; generating research reports also requires a synthesis-model API endpoint, model name, and key. Charges are determined by the services the user connects.

The project uses Python 3.10+ and a lightweight frontend, with no Node build required. PDF import uses an optional dependency. Without API configuration, users can browse the saved examples and organize material with local rules. The repository provides run instructions, model-connection configuration, tests, and release scripts. The application code is MIT-licensed.

This organization allows research outputs to participate in later work: evidence relationships enter a brief, its findings become assumptions, and new material triggers review and investigation. The research archive preserves the sources and changes throughout this process.

6. Inclusive Intelligence: an initial exploration

Jev Cowork points toward a form of intelligence worth studying further: professional judgment distributed broadly through a workflow at low cost, covering more evidence relationships and continually participating in the revision of questions. We call it Inclusive Intelligence.

First, it means expanding the scope of consideration. How much evidence an analysis can accommodate, and how many mutually constraining explanations it can recognize, depends on how limited attention is allocated. Low-cost local judgments may bring weaker signals, peripheral material, and alternative explanations into the research process earlier. Their value still requires source verification and a coherent overall argument, but more information can receive serious consideration.

It also means making research capability more accessible. Maintaining a large set of questions usually requires sustained human effort. If judgment can become a lightweight, frequent, composable operation, individuals and small teams may sustain broader research coverage at lower cost. A long-running question can continue to receive relevant changes even when nobody is actively asking about it.

These two meanings are connected through research state. An isolated judgment has limited value. Once saved as an evidence relationship, working assumption, or review condition, it can participate in future judgments. Depth enables fuller consideration; continuity preserves the results of that consideration. Together, they create research capability that accumulates over time.

This article is an initial exploration of that direction. The results reveal useful components in professional judgment, evidence organization, and continuous monitoring. They also expose the distance between local judgments and complete arguments, and between alerts and effective investigation. The next step is to observe the system in real work over longer periods: can it reduce important omissions, improve the order of investigation, and enable more people to sustain research they could not otherwise afford?

Jev Cowork provides a concrete starting point. Intuition discovers leads, reasoning develops arguments, source text imposes constraints, and people retain the final judgment. Inclusive Intelligence asks how widely this collaboration can extend: across how much information, through how many people's work, and over how much time.


References

  1. TypeSafe. Jev: models and pricing.
  2. TypeSafe. Confidence versus Probability.