Skip to content
QGI evaluation framework / LQBENCH-QGI/0.7

Generality should be measured, calibrated, demonstrated and challenged.

Public QGI R4 extends the evaluation methodology into a demonstration and evidence-publication protocol: pre-registration, cross-domain tracks, reproducible evidence packets, failures, attestations and challenges.

Capability dimensions

A QGI benchmark must test the entire quantitative cognition loop.

CROSS DOMAIN TRANSFER AND GENERALITY EVIDENCE ENABLED NO QGI CERTIFICATION
REP

Quantitative representation

Can the system identify variables, units, state, relationships and missing information needed to represent the problem quantitatively?

MOD

Model selection & construction

Can it select, construct or reject an appropriate quantitative model family under explicit assumptions?

SIM

Simulation

Can it operate deterministic, stochastic or domain-specific simulators correctly and preserve their conditions?

PRE

Prediction

Can it produce out-of-sample predictions with explicit error, calibration and failure conditions?

OPT

Optimisation

Can it search under explicit objectives and constraints without silently relaxing the problem?

UNC

Uncertainty & calibration

Can it quantify what is not known and remain calibrated as models, evidence and domains change?

EXP

Experiment design

Can it select measurements, simulations or interventions that reduce uncertainty or discriminate between hypotheses?

EVD

Evidence & reproducibility

Can a result be traced to exact inputs, methods, runtime, outputs, uncertainty and signed provenance?

MEM

Scientific memory

Can it retain hypotheses, failures, contradictions, evidence and unresolved questions without rewriting history?

XFR

Cross-domain transfer

Can the same intelligence architecture adapt to materially different quantitative systems rather than memorising one benchmark family?

AUT

Bounded autonomy

Can it close objective → experiment → evidence → decision loops under explicit budgets, guards and stop conditions?

Task-family calibration

Test the same kind of quantitative reasoning in different worlds.

TF-01State estimation under partial observation
TF-02Forecasting under regime shift
TF-03Constrained optimisation
TF-04Inverse problem reconstruction
TF-05Simulator-guided control
TF-06Experiment selection
TF-07Multifidelity transfer
TF-08Evidence and provenance audit

Task-family definitions are public evaluation design. They are not claims that any system has completed or passed them.

01

Cross-domain

A single finance, physics, engineering or biology benchmark cannot establish generality.

02

Held-out transfer

Domain, task-family and parameter-regime holdouts distinguish transfer from benchmark familiarity.

03

Evidence-bearing

Every result should preserve inputs, methods, runtime, outputs, uncertainty, resources and provenance.

04

Reproducible

A strong result should survive independent recomputation where the domain permits it.

Machine-readable contracts

Benchmark rules and demonstration evidence should be inspectable by people and systems.

LQBENCH-QGI/0.3 publishes the measurement specification, evaluation manifest, demonstration registry and evidence-packet schema without publishing fabricated results.

NO UNIVERSAL QGI SCORE

The protocol reports domain-native measurements and capability vectors. Cross-task normalisation is permitted only when its mathematical basis and reference baselines are explicit.

Calibration

Turn measurement vocabulary into executable evaluation methodology.

The calibration framework defines task families, reference baseline classes, held-out transfer regimes, resource accounting and machine-readable evaluation manifests.

MEASUREMENT BEFORE CLAIMS

Methodology is not evidence.

The framework defines how evidence should be produced without manufacturing a score, level, certification or benchmark victory.

LQBench can make QGI claims falsifiable rather than rhetorical.

Explore LQBench