Model capabilities
Typed, not generated
Three specialist heads return choices, binary judgments and ordinal ratings with option probabilities. No generated answer to parse. A multi-question request uses one forward pass per decision.
Held-out calibration
Tools for fitting temperatures and acceptance thresholds on dev data, separately from test evaluation. Calibration must be validated for your task; confidence is not a guarantee of correctness.
Test the evidence
The scientific battery includes empty- and shuffled-passage controls. They test sensitivity to evidence and expose shortcuts that accuracy alone can hide.
Research evidence Archived v0.2 study
Matched frozen and domain-adapted arms, with the same head
recipe and complete provided passages (systemone-v2).
The released fr_* checkpoints use the frozen backbone;
adaptation remains experimental.
| Evaluation | Type | Items | Frozen | Adapted |
|---|---|---|---|---|
| Scientific battery | choice | 448 | 0.970 ± 0.017 | 0.961 ± 0.011 |
| Scientific battery | noul | 1200 | 0.925 ± 0.006 | 0.913 ± 0.007 |
| Scientific battery | score | 1218 | 0.859 ± 0.075 | 0.779 ± 0.109 |
| GPQA main | choice | 441 | 0.295 ± 0.014 | 0.288 ± 0.018 |
| SciFact dev | noul | 332 | 0.483 ± 0.023 | 0.709 ± 0.020 |
| SciFact dev | choice (3-way) | 332 | 0.479 ± 0.018 | 0.444 ± 0.062 |
The internal battery uses constructed distractors and synthetic scoring labels. Its high accuracy is not a measure of general scientific correctness. Benchmark results also informed candidate selection; these are research estimates, not an untouched final test.
With the passage removed, the frozen choice head still agrees with the reference answer on 87.5% of items. Score is more sensitive to evidence changes, but sensitivity alone does not establish semantic grounding. GPQA remains a known weakness.
The "Release paper" link is the archived v0.2.2 release. The current manuscript reports this matched study with the evidence controls and per-seed analyses and is in preparation for arXiv submission. Later cleaned cohorts and ongoing QA experiments are separate results; the 441-item GPQA row above must not be mixed with the later 440-item rerun.
Get started
Install Sciev and download the frozen release heads. LLaDA-8B backbone weights are loaded separately; the example below requires CUDA and your own labeled, non-overlapping dev and test decision files.
pip install sciev
gh release download v0.2.3 --repo alrobles/sciev \
--pattern 'fr_*.pt' --dir release
python -m sciev.eval \
--ckpt release/fr_choice.pt --decision-type choice \
--r2-mode spanpool --r2-layers=-1,-9,-17,-25 --canonical-order \
--r2-temp-fit-decisions my_decisions_dev.jsonl \
--decisions-eval my_decisions_test.jsonl \
--device cuda --out eval.json
This example fits a temperature, not a deployment
acceptance policy. Keep training, calibration and evaluation data separate.
The package also provides a System-One-style /v1/systemone
interface via sciev[serve]; configure its model manifest and
run it locally or behind an authenticated proxy.
Built for scientific workflows
Candidate selection
Compare candidate answers against a supplied passage or structured state. Return an option distribution rather than a generated explanation.
Claim verification
Ask whether the supplied evidence supports a claim. Evaluate both support judgments and sensitivity to missing or conflicting evidence.
Routing & grading
Use typed choices to route a workflow, or ordinal levels to support review. New domains and rubrics require their own expert-labeled validation.
Cite Sciev
@software{robles_fernandez_sciev_2026,
author = {Robles-Fernández, Angel Luis},
title = {Sciev: Open System-One Scientific Decision Models},
url = {https://sciev.org},
version = {v0.2.2},
year = {2026}
}