Probing the capability frontier

Models trained on the hardest edges of the truth.

As capabilities improve, flaws hide in consensus. OpenAgent Lab trains models on an adversarial curriculum of questions designed to elicit maximum disagreement among the strongest systems, targeting the precise areas where the frontier can still advance.

The discriminative target

Finding the tractable signal in algorithmic dispute.

When strong systems agree on an incorrect forecast, the question is likely intractable. When they disagree, at least one system is leaving performance on the table. We isolate these high-disagreement questions to find where better reasoning yields better calibration.

I · Discriminating formats

Finer-grained outcome parameters expose latent flaws.

To efficiently probe the frontier, we prefer highly discriminating question formats. Binary thresholds lazily collapse nuanced beliefs; by converting to continuous or multiple-choice formats, we expose underlying algorithmic disagreements that drive model improvement.

  • More discriminating question formats
  • Rejecting simple binary thresholds
  • Exposing latent predictive variance
Contour terrain · resolving fine-grained algorithmic variances across a highly discriminating space
Contour terrain · resolving fine-grained algorithmic variances across a highly discriminating space
II · Locating intractability

Resolving tractable uncertainty versus intractable noise.

Better evidence gathering helps a model refine its view, but some events remain dominated by intractable noise. We evaluate our systems on questions where deep research causes top models to firmly disagree, isolating the tractable uncertainty where machine reasoning can truly improve.

  • Conditioning on deep research
  • Isolating tractable uncertainty
  • Identifying intractable noise
Glass optics · isolating the tractable algorithmic signal from intractable environmental noise
Glass optics · isolating the tractable algorithmic signal from intractable environmental noise
III · Iterative contention

Tracking algorithmic dispute across the resolution horizon.

Revisions happen continuously as the target date approaches. We track the disagreement score—the root-mean-square spread of log scores across the panel—watching how the adversarial consensus shifts and narrows as uncertainty resolves.

01 — Evidence

Updating conflicting positions.

New data reshapes the front curve, forcing top models to update their conflicting positions and re-evaluate the latent disagreement.

Evidence stage of the belief-model render
Ridgeline · evidence · frame 52
02 — Hypothesis

Exposing the precise axis of dispute.

As the curve splits into competing explanations, the variance in logarithmic scores explicitly exposes the precise axis of algorithmic dispute.

Hypothesis stage of the belief-model render
Ridgeline · hypothesis · frame 104
03 — Belief

Concentrating the probability mass.

Models continuously adjust their distributions as evidence accumulates. This concentrates the probability mass and shifts the measured disagreement.

Belief stage of the belief-model render
Ridgeline · belief · frame 166
04 — Forecast

Monitoring the panel divergence.

The structural disagreement is projected forward as a widening fan. The magnitude of this panel divergence is continuously monitored over the time horizon.

Forecast stage of the belief-model render
Ridgeline · forecast · frame 214
05 — Resolution

Scoring against empirical reality.

A truth line rises and the past distributions are scored. The initial divergence serves as proof that the uncertainty was tractable and worth chasing.

Resolution stage of the belief-model render
Ridgeline · resolution · frame 250
IV · Evaluating on the edge

Scoring questions by the disagreement they elicit.

We grade our training targets by computing the disagreement score: the variance of logarithmic scores across a panel of the strongest forecasters before the outcome resolves. A track record built on resolving these high-disagreement questions provides a rigorous signal on how to advance the capability frontier.

  • Scoring questions, not just models
  • Variance of logarithmic scores across the panel
  • Rewarding targets that expose frontier flaws
Instrument · evaluating the model on its resolution of intentionally difficult questions
Instrument · evaluating the model on its resolution of intentionally difficult questions
Verifiable discrimination

Auditing the frontier of algorithmic disagreement.

01

Pre-resolution scoring.

The generalized Jensen-Shannon divergence across the panel is computed and logged immediately after forecasts are submitted, long before the outcome is known.

02

Immutable resolution criteria.

By employing highly discriminating numerical and date formats with predefined bins, the criteria for evaluating an outcome are strictly fixed in advance.

03

Fixed panel evaluation.

The disagreement score is measured strictly against a fixed, public panel of established strong systems, preventing the score from being inflated by the errors of weaker models.