Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Spark-X2.5: observation protocols under a fixed reasoning budget

An AI-operated, reproducible case study of XHToken/Spark-X2.5-1.7B on 32 original bilingual conditional-probability prompts. The experiment separates changing the event being conditioned on from doing the fraction arithmetic.

Complete: all 32 primary cases were run and the raw evidence passed reconciliation. The fixed 2048-token, thinking-enabled greedy run delivered 7/32 correct parseable finals (21.875%); 25 were unparseable and 26 reached the token cap. This is an answer-delivery metric under the specified budget, not a claim that every unparseable output has wrong mathematics.

Read the full report and all 32 traces. The proposed secondary comparison was cancelled and has no results.

What is tested

Each problem rolls two independent fair dice. Four precisely specified rules retain different trials: at least one die equals a named value; exactly one equals it; the first die equals it; or an independently selected die equals it. The question asks the probability of a sum threshold among retained trials.

The last rule cannot be replaced by uniform sampling over pairs containing the named value. Pairs with two matching dice have two possible retaining selection choices. The exact oracle enumerates equally likely ordered pairs or triples, as appropriate. An independently derived formula checks it.

Four numeric blocks × four protocols × two languages make 32 cases. All cases, including truncated answers and mistakes, are retained. This is a small designed diagnostic, not a standard benchmark or an estimate of general mathematical ability.

Files

Reproduction

Use Python 3.12 and the versions in requirements-lock.txt, with a compatible CUDA GPU. The original machine uses one RTX 3060 Laptop GPU with 6 GB VRAM, native BF16, eager attention and batch size one. Runtime and driver details are recorded in the identity. No paid API or cloud inference was used.

For the original Windows/CUDA environment, create a Python 3.12 virtual environment and install the CUDA wheel before the remaining locked packages:

python -m venv .venv
.venv\Scripts\python -m pip install torch==2.8.0+cu128 --index-url https://download.pytorch.org/whl/cu128
.venv\Scripts\python -m pip install -r requirements-lock.txt

Activate that environment for the commands below. The lock file describes the original installation, not a cross-platform compatibility promise.

Download the official model into model/ at revision 448e61eb392c00f2c403185c5b56d5e0665bfaab. The public model can be downloaded without an authentication token; publication requires a Hugging Face account. Review the official custom code before enabling trust_remote_code=True. Weights are not included in this repository.

hf download XHToken/Spark-X2.5-1.7B --revision 448e61eb392c00f2c403185c5b56d5e0665bfaab --local-dir model

Run the following from the project root, using the documented environment:

python verify_model.py
python -m unittest -v
python -u run_study.py --data data/diagnostic-v1.json --output results/reproduction-2048 --max-new-tokens 2048
python analyze.py data/diagnostic-v1.json results/reproduction-2048

Use a fresh output directory for an independent reproduction. Resuming an existing run requires its exact recorded identity; do not mix outputs from different machines or settings. Seeded greedy decoding does not guarantee bit-identical output across hardware/runtime versions.

To reconcile the published original run with its frozen runner revision:

python validate_run.py results/primary-2048 --runner-revision 3aec462

The validator requires all 32 records and will reject incomplete runs.

Attribution and limitations

The user operating the GitHub account codeofwxz is the participant. An AI assistant designed the diagnostic, wrote and executed the scripts, reviewed the visible reasoning traces, and drafted the report on the user's local machine. The assistant is not represented as a human participant. Qualitative labels are not independent human annotations.

Original prompts are CC0-1.0; scripts use the MIT license. Official model weights and code remain under their own Apache 2.0 license. Fresh wording does not establish absence of training contamination. All scores measure one completion per case within 2048 generated tokens, with thinking enabled; they do not measure uncapped reasoning or the cancelled no-thinking arm.

About

Reproducible 32-case bilingual conditional-probability diagnostic of Spark-X2.5-1.7B, with complete raw outputs and AI-disclosed audit

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages