Much of the cost of knowledge work arises before an answer appears: identifying the facts that matter, allocating attention among competing explanations, and deciding what to investigate next. Even after a report is finished, a new disclosure, court ruling, or experimental result may invalidate one of its judgments. The research must continue.
This work depends on a kind of professional intuition: knowing what deserves attention, what warrants doubt, and what is sufficient. We call it taste. Jev delivers judgments directly as selections, scores, and probabilities. Its low cost and low latency create an opportunity to bring this intuition into more parts of the research process.
Jev Cowork operates along two dimensions: within a question, broadening the coverage of evidence and judgments about relationships; over time, maintaining assumptions, incorporating changes, and advancing investigations. We call this direction Inclusive Intelligence: allowing more information to be considered and more questions to receive affordable, sustained research attention.
1. Judgment and prediction: two capabilities of System 1
In this article, System 1 refers to model calls that directly produce bounded judgments without an explicit chain of thought. We tested the structured interface of jev-1.13.0.[1]
System 1 describes how a capability is invoked: quickly and directly, without constructing a full argument every time. It can be used for either prediction or judgment. The distinction between these capabilities comes from the task itself.
Prediction concerns outcomes: given the information available, what will happen, and with what probabilities? Judgment concerns choices: given goals, evidence, and constraints, what deserves belief, attention, or action? We test the latter through action selection, then extend it to the allocation of attention in research.
The two draw on shared knowledge but have different success criteria. Someone may correctly anticipate growing demand yet choose the wrong financing plan. Conversely, someone unable to predict market movements precisely may still recognize that a proposal violates a liquidity constraint, or identify the evidence that should be checked first. Understanding the world and choosing how to act require separate evaluation.
| Capability | Core question | Primary evaluation criteria |
|---|---|---|
| Prediction | Which outcomes will occur, and with what probabilities? | Probability error and calibration after outcomes are known |
| Judgment, tested here through action selection | What should be chosen under these goals and constraints? | Constraint satisfaction, trade-offs among objectives, and action quality |
Output length does not distinguish them either. A single probability can encode a complex prediction; a single option can encode a complex professional judgment. Taste is especially apparent when several choices all seem reasonable: which conditions cannot be violated, which exception actually applies, and whether a feasible plan deserves priority.
Judging actions in a given situation
Hard Pro contains nine cases across finance, law, and medicine. Each question includes twelve records and five candidate actions, requiring the model to integrate constraints across records, temporal relationships, and the costs of different plans. The questions and author-defined reference criteria were frozen before the runs. Models used neither retrieval nor tools. The results are shown below.

Jev matches Hunyuan 3 no_think and Cogito-32B direct, scoring one question below DS none and GLM low. Gemini 3.6 Flash scores 7/9. BERT-MNLI converts each action into a separate entailment judgment and scores 2/9. There is a clear gap between local semantic matching and professional action selection: the latter requires organizing conditions scattered across the material into a coherent whole.
Harder finance questions also expose the limits of this intuition. Jev and DS none both score 0/3, while DS max scores 3/3. In one financing question, Jev chooses a feasible but more expensive plan. Recognizing local conditions and resolving the overall trade-off are different levels of capability.
Latency reveals another difference. The median duration of Jev's primary-action requests is 0.846 seconds, compared with 1.887 seconds for DS none and 9.004 seconds for DS max. These are observations from the experimental pipeline. For local judgments that must recur frequently, subsecond responses already have practical value.
Probabilistic prediction of unknown outcomes
Beyond action selection, we also tested historical-event forecasting. On the seven financial questions for which all six configurations completed successfully, the Brier scores are as follows. Lower is better.
| Configuration | Brier score |
|---|---|
| DS V4.1 Flash max | 0.05347 |
| Gemini 3.6 Flash, thinking budget 0 | 0.09671 |
| GLM-5.3-Flash low | 0.11794 |
| DS V4.1 Flash none | 0.13233 |
| Hunyuan 3 no_think | 0.14357 |
| Jev | 0.15239 |
Jev trails the other configurations on this forecasting set. Together with the action-selection experiment, this gives a more specific capability profile: it can handle some professional constraints and action trade-offs, while its probability estimates for unknown events are comparatively weaker. The two experiments measure different aspects and cannot substitute for each other.
The same distinction applies to System 1. Quickly recognizing a situation's key structure and quickly producing calibrated probabilities require learning different mappings. Directly choosing a suitable action does not guarantee accurate probabilities over future outcomes. Even when an interface returns confidence, its calibration still needs to be evaluated separately.[2]
In another twelve-question experiment, having DS first generate a summary, a mechanistic interpretation, or an evidence audit and then passing it to Jev did not reduce mean Brier error relative to using Jev directly. What an explanation adds and whether that information improves probability estimation are separate questions.
Combining judgment and prediction
Real decisions often require both prediction and judgment. Prediction provides possible outcomes; judgment determines how to weigh gains, losses, and constraints. The same outcome probabilities can lead to different actions under different funding horizons or costs of failure. If an action changes the outcome, the prediction itself must also be conditioned on that action.
Hard Pro places rules and key conditions in the question, allowing us to observe trade-off decisions relatively directly. Open-ended research requires an earlier judgment: what information is missing, which uncertainty matters, and how much investigation is worth spending on it. Here, taste helps determine where prediction and reasoning should be applied.
This suggests a concrete role for Jev within a system. Even if it cannot yet accurately forecast a company's profit next year, it may still identify an important disclosure, notice tension between two accounts, or flag an assumption worth revisiting. The following experiments test whether these judgments can improve the research process.
2. The marginal cost of judgment
When a professional judgment becomes cheap enough, new ways of organizing research become possible.
At the experimental list price of $0.042 per million input tokens, ten thousand judgments with two thousand tokens each cost approximately $0.84 in input charges. One set of Jev Cowork research experiments made 934 independent Jev calls, using approximately 2.56 million input tokens at an estimated input cost of $0.108. Synthesis-model, retrieval, and verification costs are additional.
At this scale, we can make the objects of judgment more granular: the importance of a source, the complementarity of two pieces of evidence, the scope of a conclusion, or the effect of a new disclosure on an old assumption. These judgments can be computed separately, run in parallel, and retained as research state.
What scaling expands here is coverage across distinct judgments. Asking the same question ten times may leave the same blind spot intact; identifying ten new evidence relationships may open another path of investigation. The value of scale depends on the information gained from the additional judgments.
Knowledge work is always constrained by attention. Researchers prioritize a few urgent questions, while other uncertainties, weak signals, and long-term assumptions must wait. Low-cost judgment offers a way to broaden this attention: more material can be noticed, more connections can be proposed, and questions that nobody is actively rereading can continue to receive new information.
Inclusive Intelligence has two meanings. Within research, more evidence and explanations get a chance to be considered. Beyond an individual research task, lower maintenance costs may allow small teams and individuals to sustain attention across larger sets of questions. Jev Cowork explores these possibilities through deep research and continuous research, respectively.
3. Going deeper: broadening evidence and relationship coverage
Jev judges which materials deserve to be read together and where qualifications or counterevidence arise. Code organizes these relationships into an agenda, and a synthesis model combines it with the full source text to produce a cited report.
From finding material to building an argument
The evidence-retrieval experiment covers 380 complete documents and 90,041 source passages.

Financial target-page Hit@20 rises from 0.107 with lexical retrieval to 0.929 with Jev ranking; DS max scores 0.893 in the same round. Jev also leads on legal document ranking, but DS max still leads on positive-evidence coverage.
Finding material is not the same as completing an argument. Across six research questions on SVB and Akorn, Jev evidence composition plus DS max receives a source-verification score of 78.33/100, versus 90.83/100 for DS max given the complete evidence directly. Filtering can discard definitions and table headers. Report generation therefore preserves the full source text, letting the agenda guide reading rather than replace the sources.
Intuition proposes connections; sources constrain conclusions
For four research tasks on 3M and Waterkeeper, 512 distinct Jev relationship judgments form an agenda, which DS medium uses alongside complete source pages to write the reports.
| 3M / Waterkeeper, four tasks | Jev agenda + DS medium | DS max alone | DS none alone |
|---|---|---|---|
| Research-task success | 3/4 | 3/4 | 2/4 |
| Key facts verified | 20/20 | 20/20 | 15/20 |
| Predeclared major errors | 0 | 0 | 1 |
| Audited citations supported by sources | 60/61 | 57/59 | 38/49 |
| Writing output tokens, including reasoning | 49,993 | 67,512 | 11,137 |
| Mean writing time | 55.9 seconds | 72.4 seconds | 14.6 seconds |
Citations underwent model-assisted source verification. Token counts are totals across the four tasks; time is the mean duration of a single DS writing request. Agenda construction and research evaluation are accounted for separately.
The agenda-plus-medium configuration matches max alone on task success and key-fact verification. Each configuration has one legal report that fails delivery checks because of citation-number errors. The former uses 25.9% fewer writing output tokens and takes 22.7% less time on average. Building the four agendas with Jev costs an estimated $0.0624 at list price. These results reflect the combination of an agenda and a lower reasoning setting.
Beyond these numbers, the reports also leave questions for further investigation. At 3M, earnings per share rise slightly while adjusted earnings and cash flow weaken. Rather than stopping at "declining earnings quality," the report considers raw-material and exchange-rate effects alongside volume, pricing, and productivity. It asks how much of the decline might ease with external conditions, and how much stems from the business itself. If temporary cost pressure is the main cause, the assessment of sustained earnings may become less negative. If volume and unit profitability also continue to weaken, the operating performance of the remaining businesses needs closer examination. Follow-up research thus has concrete targets, rather than merely an instruction to "keep watching the financial reports."
In Waterkeeper, the connection comes from different pages of the same judgment. One passage says that a party did not concede the waters' relative permanence but did not directly dispute that fact either; another says that it challenged the opposing party's assertion of "permanent standing waters." Reading these passages together, the report proposes a line of inquiry: return to the pretrial statements and trial record to establish which physical facts were conceded and which remain disputed. This bears on what additional evidence is needed under the new legal standard and which questions subsequent proceedings can address.
The information gain comes from connecting existing evidence. No new source material is added, but scattered facts are organized into explanations that need to be distinguished and premises that need to be established. Here, the agenda provides research leads, not answers. The next step still requires ordering the work: which question is most likely to change the conclusion, which checks depend on earlier results, and what evidence would justify stopping.
4. Over time: keeping research active
Judgments in reports are often conditional: "the company is still prioritizing capital preservation," "the conclusion depends on a regulatory condition," or "next quarter's cash flow needs watching." These conditions are rarely maintained separately. When new information arrives, a researcher must recover the old judgment and reconnect it to the evidence.
Jev Cowork saves findings from briefs as working assumptions, records conditions for review, and continuously matches subsequent material against them. Changes worth attention generate alerts; evidence obtained through investigation enters the same research archive.
CVS: updating a capital-allocation assumption
At the end of 2020, CVS still had substantial unused share-repurchase authorization but made no repurchases in the fourth quarter. A researcher could retain a capital-allocation assumption: is the company still prioritizing capital preservation?
The subsequent 2022 annual report discloses approximately $3.5 billion in actual repurchases. The system connects this new material to the earlier assumption and proposes a review alert: repurchases have resumed, so the capital-allocation judgment deserves updating.

The basic unit of research thus expands to relationships among questions, assumptions, and evidence. A judgment left by one report can continue to participate in subsequent research.
Continuous monitoring and investigative actions
The continuous-research experiment covers eight corporate or legal topics, backed by 397 complete documents. Each case contains 256 candidate materials and 96 chronological updates. Attention allocation, continuous monitoring, and investigative actions are tested separately.

All 12 alerts delivered by Jev receive source support, compared with 7/12 for DS and 2/18 for rules. The higher alert quality makes this approach worth exploring. However, Jev's mean timely recall across cases is only 37.5%, leaving many omissions. At present, it is suited to assisting researchers with ongoing attention.
Attention allocation and investigation depend more heavily on the specific workflow design. Among the six attention-allocation cases that passed the control checks, Jev fully supports 13/24 questions, versus 12/24 for rules. In the investigation task, Jev scores 18/32, while rules reach 20/32. The Block case contains a path to complete evidence collection within budget, but Jev's actual actions sufficiently support only one of the four questions.
This exposes a deeper requirement for continuous research: the system must remember what has been resolved, what is still missing, and choose the next step accordingly. Judgments need shared state. Otherwise, even when every call is cheap, many local choices may repeatedly expend effort on the same question.
5. Jev Cowork: a collaborative research workspace
Jev Cowork is currently a research prototype centered on case demonstrations. It organizes the methods above into an interactive workspace for replaying existing results, observing research workflows, and trying new material. The current implementation is built primarily around the cases in this article; generality, stability, and the experience of long-term use still need improvement.
Project repository: Yii-Jing/Jev-Cowork.
The repository connects deep and continuous research through a unified research archive, linking material import, evidence organization, and assumption maintenance to demonstrate the basic form of this collaboration.
The workspace contains five connected views:
| View | Primary function |
|---|---|
| Deep research | Inspect evidence relationships and organize a research agenda |
| Evidence queue | Retrieve source text and track reading and verification status |
| Research brief | Read cited reports alongside their sources |
| Continuous research | Maintain working assumptions and receive updates and review alerts |
| Investigation and notes | Save follow-up leads and research notes, and export archives |
The repository includes two complete examples. The 3M example contains 136 sets of saved Jev relationship judgments, allowing comparisons between agenda-guided reports and reports generated through direct synthesis. The CVS example shows how capital-allocation assumptions connect to subsequent disclosures. Both examples can be replayed locally to inspect the correspondence among source text, judgments, and reports.
Users can import TXT, Markdown, JSON, or PDF material to create their own research archives. Materials, notes, and run records are stored locally. Archives support Markdown and JSON export, as well as restoration from JSON. Live model analysis requires the user's own Jev API key; generating research reports also requires a synthesis-model API endpoint, model name, and key. Charges are determined by the services the user connects.
The project uses Python 3.10+ and a lightweight frontend, with no Node build required. PDF import uses an optional dependency. Without API configuration, users can browse the saved examples and organize material with local rules. The repository provides run instructions, model-connection configuration, tests, and release scripts. The application code is MIT-licensed.
This organization allows research outputs to participate in later work: evidence relationships enter a brief, its findings become assumptions, and new material triggers review and investigation. The research archive preserves the sources and changes throughout this process.
6. Inclusive Intelligence: an initial exploration
Jev Cowork points toward a form of intelligence worth studying further: professional judgment distributed broadly through a workflow at low cost, covering more evidence relationships and continually participating in the revision of questions. We call it Inclusive Intelligence.
First, it means expanding the scope of consideration. How much evidence an analysis can accommodate, and how many mutually constraining explanations it can recognize, depends on how limited attention is allocated. Low-cost local judgments may bring weaker signals, peripheral material, and alternative explanations into the research process earlier. Their value still requires source verification and a coherent overall argument, but more information can receive serious consideration.
It also means making research capability more accessible. Maintaining a large set of questions usually requires sustained human effort. If judgment can become a lightweight, frequent, composable operation, individuals and small teams may sustain broader research coverage at lower cost. A long-running question can continue to receive relevant changes even when nobody is actively asking about it.
These two meanings are connected through research state. An isolated judgment has limited value. Once saved as an evidence relationship, working assumption, or review condition, it can participate in future judgments. Depth enables fuller consideration; continuity preserves the results of that consideration. Together, they create research capability that accumulates over time.
This article is an initial exploration of that direction. The results reveal useful components in professional judgment, evidence organization, and continuous monitoring. They also expose the distance between local judgments and complete arguments, and between alerts and effective investigation. The next step is to observe the system in real work over longer periods: can it reduce important omissions, improve the order of investigation, and enable more people to sustain research they could not otherwise afford?
Jev Cowork provides a concrete starting point. Intuition discovers leads, reasoning develops arguments, source text imposes constraints, and people retain the final judgment. Inclusive Intelligence asks how widely this collaboration can extend: across how much information, through how many people's work, and over how much time.
References
- TypeSafe. Jev: models and pricing.
- TypeSafe. Confidence versus Probability.