Quantitative representation
Can the system identify variables, units, state, relationships and missing information needed to represent the problem quantitatively?
The LargeQuant QGI measurement layer evaluates a vector of quantitative capabilities across materially different domains. R4 carries the calibrated measurement framework into pre-registered demonstrations and evidence publication without collapsing results into one unsupported score.
Can the system identify variables, units, state, relationships and missing information needed to represent the problem quantitatively?
Can it select, construct or reject an appropriate quantitative model family under explicit assumptions?
Can it operate deterministic, stochastic or domain-specific simulators correctly and preserve their conditions?
Can it produce out-of-sample predictions with explicit error, calibration and failure conditions?
Can it search under explicit objectives and constraints without silently relaxing the problem?
Can it quantify what is not known and remain calibrated as models, evidence and domains change?
Can it select measurements, simulations or interventions that reduce uncertainty or discriminate between hypotheses?
Can a result be traced to exact inputs, methods, runtime, outputs, uncertainty and signed provenance?
Can it retain hypotheses, failures, contradictions, evidence and unresolved questions without rewriting history?
Can the same intelligence architecture adapt to materially different quantitative systems rather than memorising one benchmark family?
Can it close objective → experiment → evidence → decision loops under explicit budgets, guards and stop conditions?
A system that dominates one domain remains specialised until transfer is demonstrated under held-out task-family, domain or regime conditions.
Reusable quantitative problem structures that can be instantiated across domains.
Naive, classical, specialised, general-purpose and domain-authority references.
Task, domain, parameter, noise, objective and tool shifts.
Compute and tool-use accounting alongside capability metrics.
task_idtask_versiontask_familydomainstate_definitionobjectiveconstraintsinput_contractallowed_models_toolsoutput_contractprimary_metricsecondary_metricsuncertainty_requirementevidence_requirementresource_budgetheld_out_regimereproducibility_procedurefailure_conditionstask_manifest_hashdataset_and_source_lineagemodel_solver_simulator_versionsruntime_environmentrandomness_and_seed_policyraw_outputsderived_metricsuncertainty_and_calibrationresource_accountingexecution_evidenceprovenance_and_signaturesR4 defines how demonstrations must be registered and how evidence must be published. Public result arrays and calibrated level thresholds remain empty until real runs exist.
QGI demonstrations