Quantitative representation
Can the system identify variables, units, state, relationships and missing information needed to represent the problem quantitatively?
Public QGI R4 extends the evaluation methodology into a demonstration and evidence-publication protocol: pre-registration, cross-domain tracks, reproducible evidence packets, failures, attestations and challenges.
Can the system identify variables, units, state, relationships and missing information needed to represent the problem quantitatively?
Can it select, construct or reject an appropriate quantitative model family under explicit assumptions?
Can it operate deterministic, stochastic or domain-specific simulators correctly and preserve their conditions?
Can it produce out-of-sample predictions with explicit error, calibration and failure conditions?
Can it search under explicit objectives and constraints without silently relaxing the problem?
Can it quantify what is not known and remain calibrated as models, evidence and domains change?
Can it select measurements, simulations or interventions that reduce uncertainty or discriminate between hypotheses?
Can a result be traced to exact inputs, methods, runtime, outputs, uncertainty and signed provenance?
Can it retain hypotheses, failures, contradictions, evidence and unresolved questions without rewriting history?
Can the same intelligence architecture adapt to materially different quantitative systems rather than memorising one benchmark family?
Can it close objective → experiment → evidence → decision loops under explicit budgets, guards and stop conditions?
Task-family definitions are public evaluation design. They are not claims that any system has completed or passed them.
A single finance, physics, engineering or biology benchmark cannot establish generality.
Domain, task-family and parameter-regime holdouts distinguish transfer from benchmark familiarity.
Every result should preserve inputs, methods, runtime, outputs, uncertainty, resources and provenance.
A strong result should survive independent recomputation where the domain permits it.
LQBENCH-QGI/0.3 publishes the measurement specification, evaluation manifest, demonstration registry and evidence-packet schema without publishing fabricated results.
The protocol reports domain-native measurements and capability vectors. Cross-task normalisation is permitted only when its mathematical basis and reference baselines are explicit.
The calibration framework defines task families, reference baseline classes, held-out transfer regimes, resource accounting and machine-readable evaluation manifests.
The framework defines how evidence should be produced without manufacturing a score, level, certification or benchmark victory.