Skip to content
QGI calibration / LQBENCH-QGI/0.7

Calibration separates capability from benchmark familiarity.

Public QGI R4 preserves the calibration rules and carries them forward into demonstration registration and evidence publication before any QGI capability threshold can be defended.

Reference baselines

A general system should not look strong merely because the comparison is weak.

NO REPORTED SCORES
B0

Naive/domain-trivial baseline

Establish whether the task is materially harder than persistence, random, mean or other domain-trivial strategies.

B1

Classical quantitative baseline

Provide a transparent statistical, numerical or optimisation reference appropriate to the task.

B2

Specialised model baseline

Measure against a model tuned for the task or domain rather than rewarding generality for weak specialised competition.

B3

Tool-using general-purpose AI baseline

Separate quantitative-native capability from language-centric systems augmented with code, calculators or domain tools.

B4

Domain authority / solver reference

Where available, anchor performance to accepted solver, simulator, laboratory or expert-computation output.

Held-out transfer regimes

Test what changes when familiarity is removed.

H1

Task-family holdout

Evaluate on tasks whose templates and instances were not used for benchmark tuning.

H2

Domain holdout

Evaluate transfer into a materially different quantitative domain.

H3

Parameter-regime shift

Move the system outside the parameter region used for calibration.

H4

Noise / missingness shift

Alter observation noise, sparsity or missing-data structure.

H5

Objective / constraint shift

Change optimisation goals or feasibility boundaries while keeping the underlying system related.

H6

Tool substitution

Swap or remove a solver, simulator or tool to test whether capability depends on one memorised execution path.

Calibration methodology

Declare the rules before observing the result.

01

Pre-register task manifests, allowed tools, metrics, budgets and failure rules before observing evaluation outcomes.

02

Use domain-native metrics with explicit directionality; do not normalise unlike metrics without a mathematically declared transformation.

03

Report repeated-run uncertainty using bootstrap or task-appropriate confidence intervals when stochasticity is material.

04

Measure probabilistic calibration separately from predictive accuracy whenever a system reports uncertainty.

05

Use paired seeds or matched instances when comparing systems so variance is not mistaken for capability.

06

Report baseline-relative gain only when both baseline and candidate were evaluated under the same task contract and resource envelope.

07

Preserve failed, timed-out, infeasible, contradictory and challenged runs in the evaluation record.

Resource-normalised evaluation

Do not confuse brute-force expenditure with quantitative intelligence.

R4 requires the evaluation envelope to record the resources used to produce a result. Resource metrics are reported alongside capability metrics rather than hidden behind a single score.

wall_clock_secondscpu_time_secondsaccelerator_time_secondspeak_memory_bytessolver_callssimulation_stepsexternal_tool_callsmodel_invocationstoken_usage_if_applicable
CALIBRATION BOUNDARY

No QGI level threshold is awarded by R4.

The protocol now defines how thresholds could be calibrated. Actual thresholds require executed reference tasks, reproducible baseline evidence and review. The public specification therefore keeps calibrated_thresholds empty.

Calibration is useful only when the evaluation is cross-domain.

See the evaluation design