Skip to content
LQBench · Quantitative reliability standard

Benchmark whether the model can be trusted with the numbers.

Evaluate mathematics, statistics, finance, science, engineering, operations, agentic quantitative workflows and formal verification under explicit execution conditions.

Five execution modes

Methodology before leaderboard.

Closed, Standard Tools, Native, Agentic and Formal modes are reported separately. Hidden and rotating test sets, contamination canaries, immutable methodology versions and signed result packets protect interpretability.

ClosedStandard ToolsNativeAgenticFormalPrivate enterprise suites
01

Explicit methodology

Every benchmark says what is measured, which direction is better and what guards apply.

02

Evidence-linked publication

Public results require the platform's publication/evidence conditions rather than an unsourced marketing number.

03

Cross-domain honesty

A molecular RMSE, physics residual and reliability probability are not forced into a universal score unless a defensible normalisation exists.

04

Institutional provenance

Registered institution keys, externally signed submissions and provenance federation extend benchmark evidence across organisational boundaries.

Comparability should be earned mathematically, not created by a dashboard.
5benchmark definitions
0opt-in public runs
LQEPevidence required to publish
lqbench.forecast.one-step.naive-baseline @ 1.0.0

One-Step Forecast Skill vs Naive

Walk-forward one-step forecast evaluation against the last-observation naive baseline. Primary score is RMSE skill: 1 - model_RMSE / naive_RMSE.

Primary metricskill_score_rmse_vs_naiveHigher is better
out-of-samplewalk-forwardno-future-observation-in-training-window
No verified public result has been published for this benchmark yet.
lqbench.probabilistic.interval-calibration @ 1.0.0

Probabilistic Interval Calibration

Rolling out-of-sample interval calibration. Primary score is absolute empirical coverage error relative to the declared nominal interval; lower is better.

Primary metriccoverage_absolute_errorLower is better
out-of-samplerolling-originseeded-bootstrapno-future-observation-in-training-window
No verified public result has been published for this benchmark yet.
lqbench.forecast.model-competition @ 1.0.0

Walk-Forward Forecast Model Competition

Cross-model out-of-sample competition among naive, AR(1) and linear-trend forecasting candidates. Primary score is the selected model RMSE skill versus the naive baseline.

Primary metricskill_score_rmse_vs_naiveHigher is better
out-of-samplewalk-forwardcross-model-comparisonstable-tie-breakno-future-observation-in-training-window
No verified public result has been published for this benchmark yet.
lqbench.streaming.prequential-forecast @ 1.0.0

Prequential Streaming Forecast Skill

Sequential one-step-ahead evaluation where every forecast is produced before its realized observation. Primary score is selected-model RMSE skill versus the naive baseline.

Primary metricskill_score_rmse_vs_naiveHigher is better
prequentialout-of-samplestreamingno-future-observation-in-training-window
No verified public result has been published for this benchmark yet.
lqbench.streaming.challenger-cycle @ 1.0.0

Autonomous Champion/Challenger Cycle

Prequential champion/challenger score with an explicit sequential drift trigger and deterministic research recommendation.

Primary metricskill_score_rmse_vs_naiveHigher is better
prequentialstreamingdrift-awarechampion-challengerno-future-observation-in-training-window
No verified public result has been published for this benchmark yet.
QGI measurement programme / LQBENCH-QGI/0.1

LQBench is extending from model benchmarks toward cross-domain quantitative-intelligence measurement.

The proposed QGI layer measures representation, modelling, simulation, prediction, optimisation, uncertainty, experiment design, evidence, memory, transfer and bounded autonomy. It publishes no QGI score or certification in R2.

PROPOSED
QGI calibration programme / LQBENCH-QGI/0.2

LQBench now has a proposed cross-domain evaluation protocol.

R3 adds task families, reference baseline classes, held-out transfer regimes, uncertainty calibration, resource accounting and evidence-bearing evaluation manifests. No QGI level or universal score is awarded.

CALIBRATE
LQBENCH-QGI/0.3

Publish the evidence behind quantitative-generalisation claims.

The QGI benchmark programme now defines demonstration manifests and evidence packets so future results can be inspected, reproduced, attested and challenged.

LQBENCH-QGI/0.4

Move from publishable protocol to verifiable result records.

R5 adds the public result bundle contract, operator-only publication tooling, evidence hashing and read-only verification endpoints.

LQBENCH-QGI/0.5

Validation is evidence, not voting.

Published QGI result records can now accumulate signed independent reproduction, methodology review, corroboration, contradictory evidence, provenance concern and limitation statements. No composite trust score is computed.

Validation protocol →
LQBENCH-QGI/0.6

From isolated evidence to comparable, federated evidence.

R7 adds same-task comparison cohorts and validation federation. It permits task-specific evidence comparison while refusing a universal QGI score, cross-task winner, reviewer reputation score or certification shortcut.

Comparative evaluation → · Validation federation →
LQBENCH-QGI/0.7

From same-task comparison to cross-domain transfer evidence.

R8 introduces LQBENCH-QGI-TRANSFER/0.1 and LQBENCH-QGI-GENERALITY/0.1. The benchmark can now publish exact cross-domain transfer relations and evidence coverage without collapsing them into a universal generality score or QGI level.

Cross-domain transfer → · Generality evidence →
Beyond leaderboards

Benchmark evidence can become institutional proof.

Benchmark evidence can extend into signed scorecards, explicit public publication, evidence export, institutional registries, cryptographic external submissions and source-preserving federation.

NO UNIVERSAL WINNER

LargeQuant does not claim that unrelated domain metrics can be collapsed into a universal ranking of models or institutions.