{"specification":"LargeQuant QGI Measurement, Calibration, Demonstration, Execution, Validation, Comparison, Federation, Transfer and Generality Specification","version":"0.7","protocol":"LQBENCH-QGI\/0.7","authority":"PUBLIC-QGI-R8","status":"cross-domain-transfer-and-generality-evidence-enabled-no-qgi-certification","calibration_status":"calibration-protocol-published-no-level-thresholds-awarded","published_date":"2026-08-31","canonical_url":"https:\/\/largequant.com\/qgi\/benchmark\/specification","machine_readable_url":"https:\/\/largequant.com\/qgi\/benchmark\/specification.json","evaluation_manifest_url":"https:\/\/largequant.com\/qgi\/benchmark\/evaluation-manifest.json","claim_boundary":["QGI is a research direction and proposed architecture.","LargeQuant does not claim to have achieved general intelligence.","Existing LargeQuant capabilities are components relevant to QGI, not proof that QGI has been achieved.","LQBENCH-QGI\/0.7 adds cross-domain transfer evidence and generality coverage over preserved result, comparison and validation registries; it does not award a QGI level.","No universal generality score, universal QGI score, cross-domain winner or QGI certification is published by Public QGI R8.","Reference baseline classes and task families are evaluation requirements, not reported benchmark results."],"measurement_doctrine":"Generality should be measured as a vector of evidence-bearing capabilities across materially different quantitative systems, with explicit task contracts, held-out transfer regimes, calibrated uncertainty, reproducibility and resource accounting.","calibration_doctrine":"Calibration must separate task difficulty, domain transfer, uncertainty quality, evidence completeness and resource use before any capability-level threshold can be defended.","dimensions":{"representation":{"code":"REP","name":"Quantitative representation","question":"Can the system identify variables, units, state, relationships and missing information needed to represent the problem quantitatively?"},"modeling":{"code":"MOD","name":"Model selection & construction","question":"Can it select, construct or reject an appropriate quantitative model family under explicit assumptions?"},"simulation":{"code":"SIM","name":"Simulation","question":"Can it operate deterministic, stochastic or domain-specific simulators correctly and preserve their conditions?"},"prediction":{"code":"PRE","name":"Prediction","question":"Can it produce out-of-sample predictions with explicit error, calibration and failure conditions?"},"optimization":{"code":"OPT","name":"Optimisation","question":"Can it search under explicit objectives and constraints without silently relaxing the problem?"},"uncertainty":{"code":"UNC","name":"Uncertainty & calibration","question":"Can it quantify what is not known and remain calibrated as models, evidence and domains change?"},"experiment_design":{"code":"EXP","name":"Experiment design","question":"Can it select measurements, simulations or interventions that reduce uncertainty or discriminate between hypotheses?"},"evidence":{"code":"EVD","name":"Evidence & reproducibility","question":"Can a result be traced to exact inputs, methods, runtime, outputs, uncertainty and signed provenance?"},"memory":{"code":"MEM","name":"Scientific memory","question":"Can it retain hypotheses, failures, contradictions, evidence and unresolved questions without rewriting history?"},"transfer":{"code":"XFR","name":"Cross-domain transfer","question":"Can the same intelligence architecture adapt to materially different quantitative systems rather than memorising one benchmark family?"},"autonomy":{"code":"AUT","name":"Bounded autonomy","question":"Can it close objective \u2192 experiment \u2192 evidence \u2192 decision loops under explicit budgets, guards and stop conditions?"}},"evaluation_domains":{"stochastic-economic":{"name":"Stochastic & economic systems","examples":["forecasting","risk","market dynamics","resource allocation"]},"physical-dynamical":{"name":"Physical & dynamical systems","examples":["state evolution","control","physics simulation","trajectory prediction"]},"engineering-design":{"name":"Engineering & design systems","examples":["constraint optimisation","response surfaces","reliability","digital twins"]},"molecular-biological":{"name":"Molecular & biological systems","examples":["property prediction","assay reasoning","candidate ranking","biological uncertainty"]},"materials-scientific":{"name":"Materials & scientific systems","examples":["composition","multifidelity simulation","experimental design","scientific discovery"]}},"levels":{"QGI-0":{"label":"Specialised","description":"Performs a defined quantitative task under a fixed problem contract."},"QGI-1":{"label":"Adaptive","description":"Adapts models or tools within one quantitative domain while preserving uncertainty and evidence."},"QGI-2":{"label":"Integrated","description":"Coordinates multiple quantitative capabilities within a discipline, including model\/tool selection and verification."},"QGI-3":{"label":"Cross-domain","description":"Transfers quantitative reasoning across materially different domains under held-out tasks and explicit evidence requirements."},"QGI-4":{"label":"Autonomous scientific","description":"Generates hypotheses, designs experiments, evaluates evidence and iterates within bounded research programmes."},"QGI-5":{"label":"Open quantitative intelligence","description":"Generalises to unfamiliar measurable systems while remaining bounded, calibrated, reproducible and evidence-bearing."}},"level_policy":["No level is awarded by Public QGI R8.","Numeric level thresholds remain intentionally unassigned until reference tasks are executed against reproducible baselines and independently reviewed.","A future level claim must report the complete dimension vector, domain coverage, held-out transfer results, uncertainty quality and resource envelope.","Cross-domain claims require held-out task families from materially different quantitative systems.","Evidence, uncertainty and reproducibility are gating requirements rather than optional bonus metrics."],"task_families":{"state-estimation":{"code":"TF-01","name":"State estimation under partial observation","tests":["representation","uncertainty","prediction"],"description":"Infer latent quantitative state from incomplete, noisy or delayed observations without inventing unobserved certainty."},"forecast-shift":{"code":"TF-02","name":"Forecasting under regime shift","tests":["prediction","uncertainty","transfer"],"description":"Predict beyond the calibration regime and identify when distribution shift invalidates prior assumptions."},"constrained-optimisation":{"code":"TF-03","name":"Constrained optimisation","tests":["modeling","optimization","evidence"],"description":"Optimise an explicit objective while respecting hard constraints, feasibility and reproducible solver conditions."},"inverse-problem":{"code":"TF-04","name":"Inverse problem reconstruction","tests":["representation","modeling","uncertainty"],"description":"Recover plausible hidden parameters or structures from observable consequences and report non-identifiability."},"simulator-control":{"code":"TF-05","name":"Simulator-guided control","tests":["simulation","optimization","autonomy"],"description":"Choose bounded actions in a dynamic simulator while preserving state, constraints and stop conditions."},"experiment-selection":{"code":"TF-06","name":"Experiment selection","tests":["experiment_design","uncertainty","memory"],"description":"Select the next measurement or simulation to reduce uncertainty or discriminate between competing hypotheses."},"multifidelity-transfer":{"code":"TF-07","name":"Multifidelity transfer","tests":["simulation","transfer","modeling"],"description":"Compose cheap and expensive quantitative engines while preserving error bounds and domain validity."},"evidence-audit":{"code":"TF-08","name":"Evidence and provenance audit","tests":["evidence","memory","autonomy"],"description":"Reconstruct how a quantitative conclusion was produced and fail closed when lineage, runtime or uncertainty evidence is missing."}},"baseline_classes":[{"id":"B0","name":"Naive\/domain-trivial baseline","purpose":"Establish whether the task is materially harder than persistence, random, mean or other domain-trivial strategies."},{"id":"B1","name":"Classical quantitative baseline","purpose":"Provide a transparent statistical, numerical or optimisation reference appropriate to the task."},{"id":"B2","name":"Specialised model baseline","purpose":"Measure against a model tuned for the task or domain rather than rewarding generality for weak specialised competition."},{"id":"B3","name":"Tool-using general-purpose AI baseline","purpose":"Separate quantitative-native capability from language-centric systems augmented with code, calculators or domain tools."},{"id":"B4","name":"Domain authority \/ solver reference","purpose":"Where available, anchor performance to accepted solver, simulator, laboratory or expert-computation output."}],"held_out_regimes":[{"id":"H1","name":"Task-family holdout","description":"Evaluate on tasks whose templates and instances were not used for benchmark tuning."},{"id":"H2","name":"Domain holdout","description":"Evaluate transfer into a materially different quantitative domain."},{"id":"H3","name":"Parameter-regime shift","description":"Move the system outside the parameter region used for calibration."},{"id":"H4","name":"Noise \/ missingness shift","description":"Alter observation noise, sparsity or missing-data structure."},{"id":"H5","name":"Objective \/ constraint shift","description":"Change optimisation goals or feasibility boundaries while keeping the underlying system related."},{"id":"H6","name":"Tool substitution","description":"Swap or remove a solver, simulator or tool to test whether capability depends on one memorised execution path."}],"calibration_methods":["Pre-register task manifests, allowed tools, metrics, budgets and failure rules before observing evaluation outcomes.","Use domain-native metrics with explicit directionality; do not normalise unlike metrics without a mathematically declared transformation.","Report repeated-run uncertainty using bootstrap or task-appropriate confidence intervals when stochasticity is material.","Measure probabilistic calibration separately from predictive accuracy whenever a system reports uncertainty.","Use paired seeds or matched instances when comparing systems so variance is not mistaken for capability.","Report baseline-relative gain only when both baseline and candidate were evaluated under the same task contract and resource envelope.","Preserve failed, timed-out, infeasible, contradictory and challenged runs in the evaluation record."],"resource_accounting":["wall_clock_seconds","cpu_time_seconds","accelerator_time_seconds","peak_memory_bytes","solver_calls","simulation_steps","external_tool_calls","model_invocations","token_usage_if_applicable"],"task_contract":["task_id","task_version","task_family","domain","state_definition","objective","constraints","input_contract","allowed_models_tools","output_contract","primary_metric","secondary_metrics","uncertainty_requirement","evidence_requirement","resource_budget","held_out_regime","reproducibility_procedure","failure_conditions"],"evidence_contract":["task_manifest_hash","dataset_and_source_lineage","model_solver_simulator_versions","runtime_environment","randomness_and_seed_policy","raw_outputs","derived_metrics","uncertainty_and_calibration","resource_accounting","execution_evidence","provenance_and_signatures"],"evaluation_manifest":["evaluation_id","protocol_version","system_id","system_version","task_manifest_hash","baseline_ids","held_out_regime","run_count","resource_budget","environment_hash","evidence_bundle_id","started_at","completed_at"],"result_record":["evaluation_id","task_id","task_version","domain","dimension_vector","primary_metric","secondary_metrics","uncertainty_metrics","baseline_relative_metrics","resource_metrics","status","failure_reason","evidence_hash"],"demonstration_status":"protocol-published-no-public-qgi-results","demonstration_doctrine":"A QGI demonstration must be pre-registered, execute under a declared system and resource envelope, publish reproducible evidence, preserve failures and challenges, and separate demonstration evidence from any later capability-level claim.","demonstration_principles":["Register the demonstration objective, domain, task families, held-out regimes, metrics, baselines and resource budget before execution.","Use the same declared intelligence architecture across the demonstration; domain-specific tools may vary only when the task contract permits them.","Publish successful, failed, infeasible, contradictory, timed-out and challenged runs rather than curating only favorable outcomes.","Bind every public result to exact task, system, environment, dataset, model, solver, simulator and evidence hashes.","Keep benchmark methodology separate from evidence, and evidence separate from any future QGI level or certification claim.","Treat cross-domain transfer as evidence only when the held-out regime and domain separation were declared before results were observed."],"demonstration_tracks":{"human-systems":{"code":"DT-01","name":"Human systems \/ HumanTwin","domain":"molecular-biological","focus":"state estimation, predictive simulation, uncertainty and evidence across measurable human-system variables","status":"reference-track-no-public-result"},"mission-engineering":{"code":"DT-02","name":"Mission & engineering \/ Stratomission","domain":"engineering-design","focus":"dynamics, simulation, constrained optimisation, control and reproducible engineering evidence","status":"reference-track-no-public-result"},"financial-systems":{"code":"DT-03","name":"Stochastic & financial systems","domain":"stochastic-economic","focus":"forecasting, risk, optimisation, regime shift and calibrated uncertainty","status":"reference-track-no-public-result"},"scientific-discovery":{"code":"DT-04","name":"Scientific discovery systems","domain":"materials-scientific","focus":"multifidelity modelling, candidate search, experiment selection, reproducibility and scientific memory","status":"reference-track-no-public-result"}},"demonstration_contract":["demonstration_id","demonstration_version","title","system_id","system_version","domain","task_family_ids","held_out_regimes","objective","constraints","input_lineage","baseline_ids","allowed_tools","resource_budget","pre_registered_metrics","uncertainty_requirement","evidence_requirement","evaluation_manifest_id","status"],"evidence_packet_contract":["packet_id","protocol_version","demonstration_id","evaluation_id","task_manifest_hash","system_manifest_hash","environment_hash","dataset_lineage_hash","model_solver_simulator_manifest","raw_output_hashes","metric_records","uncertainty_records","resource_records","failure_records","provenance_records","attestation_refs","challenge_refs","signatures","published_at"],"publication_states":["planned","registered","running","evidence-complete","published","challenged","withdrawn"],"result_publication_policy":["No result may be described as a QGI capability claim unless the corresponding task and evaluation manifests were registered before result publication.","Result pages must expose domain-native metrics, uncertainty, resource use, baselines, failures and evidence references together.","A demonstration may publish evidence without receiving any QGI level.","Challenges and contradictory evidence remain visible and must not be silently deleted from the publication history.","Public QGI R5 ships with no fabricated demonstration results, calibrated thresholds or certifications. Real result records may appear only after controlled operator publication of evidence bundles."],"demonstration_registry_url":"https:\/\/largequant.com\/qgi\/demonstrations\/registry.json","evidence_schema_url":"https:\/\/largequant.com\/qgi\/evidence\/schema.json","execution_status":"operator-controlled-public-result-registry-preserved","execution_doctrine":"R5 allows real demonstration evidence bundles to be published through an operator-only CLI path into shared runtime storage. Public web routes are read-only, record hashes are reverified on every read, evidence packets are hash-bound, and publication never implies a QGI level or certification.","result_registry_url":"https:\/\/largequant.com\/qgi\/results\/registry.json","result_bundle_schema_url":"https:\/\/largequant.com\/qgi\/results\/schema.json","execution_registry_policy":["Web publication is read-only. Result writes are accepted only through the operator CLI publisher.","Every result record is SHA-256 bound to canonical JSON and reverified before public display.","Every public result is bound to an exact evidence packet SHA-256; evidence is served only if the stored packet matches that hash.","Duplicate result identifiers are idempotent only when the canonical record hash is identical; conflicting overwrites fail closed.","R5 accepts completed, failed, infeasible, timed-out and contradictory outcomes so negative evidence is not silently removed.","A result record must set claim_scope=demonstration-evidence and qgi_level\/certification to null.","A published demonstration result is evidence, not a QGI capability-level award."],"comparison_status":"operator-controlled-same-task-comparison-registry-enabled","comparison_doctrine":"R7 compares only integrity-verified published results bound to the same task manifest. Domain-native metrics, uncertainty, resource use, failures and validation evidence remain separate; no cross-task universal score is computed.","comparison_protocol":"LQBENCH-QGI-COMPARE\/0.1","comparison_registry_url":"https:\/\/largequant.com\/qgi\/comparisons\/registry.json","comparison_protocol_url":"https:\/\/largequant.com\/qgi\/comparisons\/protocol.json","comparison_contract":["comparison_id","comparison_version","title","task_manifest_sha256","domain","held_out_regime","metric_contract","resource_policy","result_ids","claim_scope","qgi_level","certification","published_at"],"comparison_registry_policy":["Comparison writes are accepted only through the operator CLI publisher; public web routes are read-only.","Every comparison must contain at least two unique, integrity-verified published result IDs.","Every member is bound to the exact result record SHA-256 and evidence packet SHA-256 current at comparison publication.","All members must bind to the same evidence task_manifest_hash; cross-task results cannot be placed in one comparison cohort.","The comparison preserves raw primary metrics, uncertainty, resources, failures and outcome state rather than collapsing them into a universal score.","Metric direction and interpretation must be declared explicitly in metric_contract; unlike metrics are not silently normalised.","Task-specific comparisons may expose evidence-bearing differences, but they do not award a QGI level, universal winner or certification."],"federation_status":"validation-evidence-federation-enabled-no-consensus-score","federation_doctrine":"R7 federates signed validation evidence across results and comparison cohorts while preserving every statement and signer key fingerprint. Federation exposes evidence coverage and contradiction; it does not convert statement counts into scientific truth, reviewer reputation or certification.","federation_protocol":"LQBENCH-QGI-FEDERATE\/0.1","federation_protocol_url":"https:\/\/largequant.com\/qgi\/federation\/protocol.json","federation_matrix_url":"https:\/\/largequant.com\/qgi\/federation\/matrix.json","federation_policy":["Federation reads the preserved R6 signed validation registry; it does not rewrite or re-sign independent statements.","Signed statements remain bound to exact result and evidence hashes and retain their original supportive or critical type.","The same signer key appearing across results is key continuity only; it is not proof of legal identity, accreditation or reputation.","Evidence coverage matrices count published evidence and validation coverage, not intelligence, truth or trust.","Contradictory evidence remains visible and is never cancelled by a majority of supportive statements.","No consensus score, reviewer score, universal QGI score, QGI level or certification is produced by federation."],"transfer_status":"operator-controlled-cross-domain-transfer-registry-enabled","transfer_doctrine":"R8 treats transfer as an evidence-bearing relation between two integrity-verified published results from materially different quantitative domains. Every relation binds exact result, evidence, task and system-manifest hashes plus a pre-registered transfer manifest, held-out regime and adaptation budget.","transfer_protocol":"LQBENCH-QGI-TRANSFER\/0.1","transfer_registry_url":"https:\/\/largequant.com\/qgi\/transfers\/registry.json","transfer_protocol_url":"https:\/\/largequant.com\/qgi\/transfers\/protocol.json","transfer_contract":["transfer_id","transfer_version","title","source_result_id","target_result_id","source_domain","target_domain","held_out_regime","transfer_regime","transfer_manifest_sha256","adaptation_budget","adaptation_manifest_sha256","metric_contract","claim_scope","qgi_level","certification","published_at"],"transfer_registry_policy":["Transfer writes are accepted only through the operator CLI publisher; public web routes are read-only.","Every transfer binds two already-published, integrity-verified results and their exact result\/evidence hashes.","Source and target domains must be materially different and must match the domains declared by the bound result records.","Both results must belong to the same system_id lineage; zero-shot and instruction-only regimes additionally require an unchanged system_version.","Every transfer publishes an exact source task manifest hash, target task manifest hash and pre-registered transfer_manifest_sha256.","Zero-shot and instruction-only regimes must declare training_examples=0 and parameter_updates=false; adaptation_manifest_sha256 must remain null.","Few-shot and bounded-adaptation regimes must publish an exact adaptation_manifest_sha256 so system changes are not hidden.","Cross-domain transfer coverage is evidence under declared conditions; it is not a universal generality score, QGI level or certification.","Failed, infeasible, timed-out and contradictory target outcomes remain publishable evidence and are not silently removed."],"generality_status":"cross-domain-evidence-coverage-enabled-no-generality-score","generality_doctrine":"R8 exposes which systems have been tested across domain pairs and under which transfer regimes. Coverage and validation metadata remain descriptive evidence; they are not collapsed into a universal generality score, QGI level or certification.","generality_protocol":"LQBENCH-QGI-GENERALITY\/0.1","generality_protocol_url":"https:\/\/largequant.com\/qgi\/generality\/protocol.json","generality_matrix_url":"https:\/\/largequant.com\/qgi\/generality\/matrix.json","generality_policy":["Generality evidence is represented as domain-pair transfer coverage and per-system transfer graphs rather than a single score.","Counts of transfers, validations or signer keys are coverage metadata; they do not establish scientific truth or intelligence level.","A generality claim must preserve transfer regime, held-out regime, adaptation budget, uncertainty, resource use and failure evidence.","Cross-domain evidence must be inspectable down to exact result, evidence, task, system and transfer-manifest hashes.","No QGI level or certification is inferred from domain coverage alone."],"published_transfers":[],"scoring_rules":["Report a capability vector by dimension and domain; do not publish an unsupported universal QGI score.","Use domain-native primary metrics and metric direction.","Normalise across tasks only where the mathematical transformation and reference baseline are explicit.","Publish resource budgets so brute-force scale is not confused with reasoning quality.","Treat missing uncertainty or missing evidence as an evaluation failure where those requirements apply.","Use held-out regimes to distinguish transfer from benchmark familiarity.","Keep failed, contradictory and challenged results visible in the evidence history."],"calibrated_thresholds":[],"public_demonstrations":[],"results":[],"certifications":[]}