Benchmark whether the model can be trusted with the numbers.
Evaluate mathematics, statistics, finance, science, engineering, operations, agentic quantitative workflows and formal verification under explicit execution conditions.
Methodology before leaderboard.
Closed, Standard Tools, Native, Agentic and Formal modes are reported separately. Hidden and rotating test sets, contamination canaries, immutable methodology versions and signed result packets protect interpretability.
Explicit methodology
Every benchmark says what is measured, which direction is better and what guards apply.
Evidence-linked publication
Public results require the platform's publication/evidence conditions rather than an unsourced marketing number.
Cross-domain honesty
A molecular RMSE, physics residual and reliability probability are not forced into a universal score unless a defensible normalisation exists.
Institutional provenance
Registered institution keys, externally signed submissions and provenance federation extend benchmark evidence across organisational boundaries.
Comparability should be earned mathematically, not created by a dashboard.
One-Step Forecast Skill vs Naive
Walk-forward one-step forecast evaluation against the last-observation naive baseline. Primary score is RMSE skill: 1 - model_RMSE / naive_RMSE.
Probabilistic Interval Calibration
Rolling out-of-sample interval calibration. Primary score is absolute empirical coverage error relative to the declared nominal interval; lower is better.
Walk-Forward Forecast Model Competition
Cross-model out-of-sample competition among naive, AR(1) and linear-trend forecasting candidates. Primary score is the selected model RMSE skill versus the naive baseline.
Prequential Streaming Forecast Skill
Sequential one-step-ahead evaluation where every forecast is produced before its realized observation. Primary score is selected-model RMSE skill versus the naive baseline.
Autonomous Champion/Challenger Cycle
Prequential champion/challenger score with an explicit sequential drift trigger and deterministic research recommendation.
LQBench is extending from model benchmarks toward cross-domain quantitative-intelligence measurement.
The proposed QGI layer measures representation, modelling, simulation, prediction, optimisation, uncertainty, experiment design, evidence, memory, transfer and bounded autonomy. It publishes no QGI score or certification in R2.
LQBench now has a proposed cross-domain evaluation protocol.
R3 adds task families, reference baseline classes, held-out transfer regimes, uncertainty calibration, resource accounting and evidence-bearing evaluation manifests. No QGI level or universal score is awarded.
Publish the evidence behind quantitative-generalisation claims.
The QGI benchmark programme now defines demonstration manifests and evidence packets so future results can be inspected, reproduced, attested and challenged.
Move from publishable protocol to verifiable result records.
R5 adds the public result bundle contract, operator-only publication tooling, evidence hashing and read-only verification endpoints.
Validation is evidence, not voting.
Published QGI result records can now accumulate signed independent reproduction, methodology review, corroboration, contradictory evidence, provenance concern and limitation statements. No composite trust score is computed.
Validation protocol →From isolated evidence to comparable, federated evidence.
R7 adds same-task comparison cohorts and validation federation. It permits task-specific evidence comparison while refusing a universal QGI score, cross-task winner, reviewer reputation score or certification shortcut.
Comparative evaluation → · Validation federation →From same-task comparison to cross-domain transfer evidence.
R8 introduces LQBENCH-QGI-TRANSFER/0.1 and LQBENCH-QGI-GENERALITY/0.1. The benchmark can now publish exact cross-domain transfer relations and evidence coverage without collapsing them into a universal generality score or QGI level.
Cross-domain transfer → · Generality evidence →Benchmark evidence can become institutional proof.
Benchmark evidence can extend into signed scorecards, explicit public publication, evidence export, institutional registries, cryptographic external submissions and source-preserving federation.
LargeQuant does not claim that unrelated domain metrics can be collapsed into a universal ranking of models or institutions.