An AI-operated, reproducible case study of XHToken/Spark-X2.5-1.7B on 32
original bilingual conditional-probability prompts. The experiment separates
changing the event being conditioned on from doing the fraction arithmetic.
Complete: all 32 primary cases were run and the raw evidence passed reconciliation. The fixed 2048-token, thinking-enabled greedy run delivered 7/32 correct parseable finals (21.875%); 25 were unparseable and 26 reached the token cap. This is an answer-delivery metric under the specified budget, not a claim that every unparseable output has wrong mathematics.
Read the full report and all 32 traces. The proposed secondary comparison was cancelled and has no results.
Each problem rolls two independent fair dice. Four precisely specified rules retain different trials: at least one die equals a named value; exactly one equals it; the first die equals it; or an independently selected die equals it. The question asks the probability of a sum threshold among retained trials.
The last rule cannot be replaced by uniform sampling over pairs containing the named value. Pairs with two matching dice have two possible retaining selection choices. The exact oracle enumerates equally likely ordered pairs or triples, as appropriate. An independently derived formula checks it.
Four numeric blocks × four protocols × two languages make 32 cases. All cases, including truncated answers and mistakes, are retained. This is a small designed diagnostic, not a standard benchmark or an estimate of general mathematical ability.
- METHODS.md: questions, exact oracle, extraction and limitations.
- RUN.md: settings, pilot repair, interruptions and commands.
- DESIGN.md: design recorded before scored inference.
- data/diagnostic-v1.json: all prompts and oracle counts.
- results/primary-2048/identity.json: frozen run identity.
- results/primary-2048/raw.jsonl: complete saved generations, including token IDs and timings; one JSON object per case.
- results/reasoning-audit.json: AI-performed qualitative review, separate from delivered-answer scoring.
study.py,run_study.py,runlog.py,analyze.py,validate_run.py: oracle, inference, resume checks, scoring and evidence reconciliation.
Use Python 3.12 and the versions in requirements-lock.txt, with a compatible
CUDA GPU. The original machine uses one RTX 3060 Laptop GPU with 6 GB VRAM,
native BF16, eager attention and batch size one. Runtime and driver details
are recorded in the identity. No paid API or cloud inference was used.
For the original Windows/CUDA environment, create a Python 3.12 virtual environment and install the CUDA wheel before the remaining locked packages:
python -m venv .venv
.venv\Scripts\python -m pip install torch==2.8.0+cu128 --index-url https://download.pytorch.org/whl/cu128
.venv\Scripts\python -m pip install -r requirements-lock.txt
Activate that environment for the commands below. The lock file describes the original installation, not a cross-platform compatibility promise.
Download the official model into model/ at revision
448e61eb392c00f2c403185c5b56d5e0665bfaab. The public model can be downloaded
without an authentication token; publication requires a Hugging Face account.
Review the official custom code before enabling trust_remote_code=True.
Weights are not included in this repository.
hf download XHToken/Spark-X2.5-1.7B --revision 448e61eb392c00f2c403185c5b56d5e0665bfaab --local-dir model
Run the following from the project root, using the documented environment:
python verify_model.py
python -m unittest -v
python -u run_study.py --data data/diagnostic-v1.json --output results/reproduction-2048 --max-new-tokens 2048
python analyze.py data/diagnostic-v1.json results/reproduction-2048
Use a fresh output directory for an independent reproduction. Resuming an existing run requires its exact recorded identity; do not mix outputs from different machines or settings. Seeded greedy decoding does not guarantee bit-identical output across hardware/runtime versions.
To reconcile the published original run with its frozen runner revision:
python validate_run.py results/primary-2048 --runner-revision 3aec462
The validator requires all 32 records and will reject incomplete runs.
The user operating the GitHub account codeofwxz is the participant. An AI
assistant designed the diagnostic, wrote and executed the scripts, reviewed
the visible reasoning traces, and drafted the report on the user's local
machine. The assistant is not represented as a human participant. Qualitative
labels are not independent human annotations.
Original prompts are CC0-1.0; scripts use the MIT license. Official model weights and code remain under their own Apache 2.0 license. Fresh wording does not establish absence of training contamination. All scores measure one completion per case within 2048 generated tokens, with thinking enabled; they do not measure uncapped reasoning or the cancelled no-thinking arm.