Harness-1

Harness-1 · Launch Video Breakdown: Hook, Pacing & Motion Design

A 20B search agent trained with a state-externalizing harness.

AI AgentsLaunchJune 6, 2026@patpcj
0:00 · The Hook · Introducing Harness-1: A New Paradigm for AI Search Agents
0:00 / 0:00

Scene-by-scene timeline & spoken transcript

  1. The Hook

    Introducing Harness-1: A New Paradigm for AI Search Agents

    “(No spoken dialogue — Ambient electronic score with pizzicato synths and sub-bass, building energy with a riser and culminating in a sub-bass drop.)”

    On screen
    HARNESS-1 REINFORCEMENT LEARNING · STATE-EXTERNALIZING HARNESS Harness-1 A 20-billion-parameter search agent that matches the frontier at finding evidence. 20B parameters · 8 difficult benchmarks · 0.730 avg recall Fully supported by Chroma · Model trained with Tinker
    Camera
    Static, wide shot of a digital slide presentation.
    Motion
    Text reveals with subtle fades and scale animations, icon animations, and progress bar fills.
  2. Problem Agitation

    The Dual Burden of Traditional Search Agents

    “(No spoken dialogue — The electronic score continues, maintaining a steady rhythm with the 808-style drum beat. Subtle sound effects accompany on-screen text reveals.)”

    On screen
    HARNESS-1 THE PROBLEM A search agent is asked to do two jobs at once. 01 Decide what to search for. 02 Remember everything it has seen. As the transcript grows, the model spends its attention rebuilding its own memory every turn. For RL, the signal becomes poorly conditioned.
    Camera
    Static, wide shot of a digital slide presentation.
    Motion
    Sequential text reveals, icon animations (e.g., search icon appearing), and line drawing animations to illustrate connections.
  3. Product Reveal

    Harness-1: Splitting the Cognitive Load

    “(No spoken dialogue — The music gains intensity with a high-energy audio transition, risers, and sweeps, leading into a more complex and driving section of the score.)”

    On screen
    HARNESS-1 THE IDEA Split the two jobs. POLICY the decisions what to search · what to keep what to verify · when to stop HARNESS the state candidate pools · curated evidence · links · memory · budget The paper calls it stateful cognitive offloading. HARNESS-1 THE HARNESS + WORKING MEMORY Search state, kept by the environment. Candidate pool every document retrieved so far Curated set kept docs, tagged by importance · cap 30 Evidence graph entities bridging documents Verification claims checked against source text Compression observations deduplicated and shrunk Budget render kept within the context limit the model keeps only the decisions → search keep verify stop
    Camera
    Static, wide shot of a digital slide presentation.
    Motion
    Split-screen conceptualization with animated icons and text, followed by a grid layout reveal with individual card animations and icon fills. Animated flow diagrams with connecting lines and highlighted elements.
  4. Feature Teaser

    The Recipe for Trainable Search and Unseen Environment Transfer

    “(No spoken dialogue — The electronic score maintains its driving energy, with a final build-up and impactful sub-bass drop to conclude the video.)”

    On screen
    HARNESS-1 HOW IT WORKS · ONE TURN The policy decides. The harness remembers. action — curate( add, importance:high ) HARNESS environment side working memory Candidate pool all retrieved Curated set cap 30 · evict lowest Evidence graph bridges + hops Verification yes / no on claim Compression top-4 BRCS · dedup 10 Budget render 88,720 tokens POLICY 20B model reasons + decides emit 1 action / turn observation — working memory + recent turns (state, action) → (state', observation) — not just action → observation. Free-form reasoning stays in the model; bookkeeping moves to the harness. Retrieve fan_out · search · grep Inspect read · review Curate add / remove · tag Verify claim End submit set HARNESS-1 HOW IT WORKS · WHAT THE MODEL SEES Not a transcript — a structured state. == Working Memory · summarizing turns 0-12 == Query "Which Brussels synagogue, completed in 1878, was designed by Désiré De Keyser?" Curated Set [14 / 30] very_high [+] 22816: Grande Synagogue · ✓ verified high [+] 91442: Désiré De Keyser, architect high [+] 62390: Brussels synagogue fair [*] 88114: 19th-c Brussels architecture low [ ] 30119: Belgian heritage list · evicts first at 30 Document Pool — 31 total · 17 uncurated [ ] 99012: synagogues of Belgium [ ] 50441: De Keyser family records Earlier uncurated (15): 7712, 6655, 6201, 7457 (+11) [Evidence Graph] bridges Brussels — 22816, 62390, 88114, 91442 1878 — 22816, 91442, 62390 De Keyser — 22816, 91442 bridge docs: 22816, 91442 5 singletons — hops Verification claim: 1878 — De Keyser — Brussels synagogue 4 22816 yes — states 1878, De Keyser 4 62390 no — register, architect unnamed Search History T9 fan_out_search — 11 new · +6 curated [TIP] use grep_corpus for exact names T10 grep_corpus "De Keyser" — 3 hits T11 verify — 1 yes, 1 no Six kinds of state — extracted, ranked, verified, deduped, budgeted — re-derived for the model every single turn. None of it belongs in the policy's head. HARNESS-1 THE RECIPE The harness makes search trainable. 01 Warm-started curation The first good search seeds the set with its top 8 — the model is always editing, never starting from a blank page. 02 Compact state Importance tags, the evidence graph, verification records — all rendered small enough to fit. 03 Diversity incentives Reward a rhythm, not just discovery: search → curate → review → verify. A short supervised phase teaches the interface. Reinforcement learning teaches the decisions. HARNESS-1 THE RESULT · AVERAGE EVIDENCE RECALL, 8 DIFFICULT BENCHMARKS (%) 76.4 73.0 78.9 68.8 64.7 61.6 60.3 49.6 28.9 26.2 Opus-4.6 Harness-1 GPT-5.4 Sonnet-4.6 Kimi-K2.5 Tongyi DR Context-1 GPT-OSS Search-R1 GPT-OSS-20B frontier 20B frontier frontier 30B 20B 120B 32B OUR BASE 20B parameters — above the previous best open agent by +11.4 recall, and competitive with much larger frontier searchers. GPT-OSS-20B is Harness-1's own base model — the same 20B weights, untrained. The harness and RL add +46.8 recall. HARNESS-1 THE DATA Trained on about 4k examples. 4,352 221,328 HARNESS-1 A TYPICAL SEARCH AGENT Most of the behavior lives in the interface, not in the weights. HARNESS-1 TRANSFER · UNSEEN ENVIRONMENTS The biggest gains appear where it never trained. Source-family benchmarks +7.9 Held-out transfer benchmarks +17.0 2.2x larger improvement on unseen tasks than Context-1 It didn't memorize domains. It learned a reusable search workflow: plug in your corpus, retriever, and verifier — the trained policy can operate through the same harness interface. HARNESS-1 THE LESSON The interface is the method. Don't just train a bigger brain on a thin interface. Shape the interface itself — move the bookkeeping out, and the learned search behavior can transfer to new tasks and environments. HARNESS-1 small model 20B stateful harness WORKING MEMORY frontier search 0.730 RECALL Harness-1 arXiv 2606.02373 github.com/pat-jj/harness-1 huggingface.co/pat-jj/harness-1 This work is fully supported by Chroma Model trained with Tinker
    Camera
    Static, wide shot of a digital slide presentation.
    Motion
    Complex data visualizations with animated bar charts and numerical counters, highlighting key metrics. Animated progress bars and text highlights emphasize performance gains. Final summary slide with icon-based equation and fading text reveals.

Related AI Agents Product Launches

Explore all AI Agents launches →
OpenAI
Hook 9.2137.9M
OpenAIAI Agents

To ensure that artificial general intelligence benefits all of humanity

@OpenAI
Gopuff
Hook 9.277.6M
GopuffAI Agents

Gopuff introduces Go, an AI shopping assistant built with SpaceXAI: say what you need and the order is placed.

@gopuff
Elon Musk
Hook 9.263.3M
Elon MuskAI Agents

Iliad (Troy) trailer made by Grok Imagine 1.5, which was just released

@elonmusk