Naive/domain-trivial baseline
Establish whether the task is materially harder than persistence, random, mean or other domain-trivial strategies.
Public QGI R4 preserves the calibration rules and carries them forward into demonstration registration and evidence publication before any QGI capability threshold can be defended.
Establish whether the task is materially harder than persistence, random, mean or other domain-trivial strategies.
Provide a transparent statistical, numerical or optimisation reference appropriate to the task.
Measure against a model tuned for the task or domain rather than rewarding generality for weak specialised competition.
Separate quantitative-native capability from language-centric systems augmented with code, calculators or domain tools.
Where available, anchor performance to accepted solver, simulator, laboratory or expert-computation output.
Evaluate on tasks whose templates and instances were not used for benchmark tuning.
Evaluate transfer into a materially different quantitative domain.
Move the system outside the parameter region used for calibration.
Alter observation noise, sparsity or missing-data structure.
Change optimisation goals or feasibility boundaries while keeping the underlying system related.
Swap or remove a solver, simulator or tool to test whether capability depends on one memorised execution path.
Pre-register task manifests, allowed tools, metrics, budgets and failure rules before observing evaluation outcomes.
Use domain-native metrics with explicit directionality; do not normalise unlike metrics without a mathematically declared transformation.
Report repeated-run uncertainty using bootstrap or task-appropriate confidence intervals when stochasticity is material.
Measure probabilistic calibration separately from predictive accuracy whenever a system reports uncertainty.
Use paired seeds or matched instances when comparing systems so variance is not mistaken for capability.
Report baseline-relative gain only when both baseline and candidate were evaluated under the same task contract and resource envelope.
Preserve failed, timed-out, infeasible, contradictory and challenged runs in the evaluation record.
R4 requires the evaluation envelope to record the resources used to produce a result. Resource metrics are reported alongside capability metrics rather than hidden behind a single score.
wall_clock_secondscpu_time_secondsaccelerator_time_secondspeak_memory_bytessolver_callssimulation_stepsexternal_tool_callsmodel_invocationstoken_usage_if_applicableThe protocol now defines how thresholds could be calibrated. Actual thresholds require executed reference tasks, reproducible baseline evidence and review. The public specification therefore keeps calibrated_thresholds empty.