Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
127 changes: 127 additions & 0 deletions docs/ROADMAP_v0.2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
# NLAttack v0.2 roadmap (LOCAL DRAFT — not yet pushed)

v0.2.0 is the first minor version after v0.1. Primary goals: expanded domain
coverage, harness correctness fixes, and cleaner multi-version result management.
Content scope is additive — v0.1 benchmark content is untouched.

---

## 1. Harness fixes (from ultrareview 2026-06-12, confirmed ×2-of-3 Opus lenses)

These fix code bugs, not benchmark design. All required before any new results
under the v0.2 label.

| # | File | Issue | Fix |
|---|---|---|---|
| H1 | `emergence.py:63` | `hash(w) % dim` uses PYTHONHASHSEED — Tier-1 nondeterministic across processes | Replace with `mmh3.hash(w) % dim` or `int(hashlib.md5(w.encode()).hexdigest(),16) % dim` |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using 'hashlib.md5' can be relatively slow if executed in a tight loop, and 'mmh3' requires an external dependency. A faster, deterministic, and dependency-free alternative is to use 'zlib.adler32' or 'binascii.crc32' from the Python standard library (e.g., 'zlib.adler32(w.encode()) % dim').

Suggested change
| H1 | `emergence.py:63` | `hash(w) % dim` uses PYTHONHASHSEED — Tier-1 nondeterministic across processes | Replace with `mmh3.hash(w) % dim` or `int(hashlib.md5(w.encode()).hexdigest(),16) % dim` |
| H1 | emergence.py:63 | hash(w) % dim uses PYTHONHASHSEED — Tier-1 nondeterministic across processes | Replace with zlib.adler32(w.encode()) % dim or binascii.crc32(w.encode()) % dim |

| H2 | `adapters.py:193-196` | Schema-drift 200 → BOS noise / empty string; no JSON parse guard | Fail loudly on missing `prompt_length`/`results`; validate before use |
| H3 | `emergence_dashboard.py:117` | `json.dumps` with `allow_nan=True` writes invalid JSON | `allow_nan=False`; replace NaN with `null` |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In Python's json module, setting allow_nan=False in json.dumps does not automatically replace NaN values with null. Instead, it raises a ValueError when encountering any NaN (or inf, -inf) values during serialization.

To safely serialize NaN values as null in the output JSON, the data must be pre-processed to replace float NaN values with None (which serializes to null), or a custom JSON encoder must be used.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Note that calling 'json.dumps' with 'allow_nan=False' will raise a 'ValueError' if a 'NaN' value is encountered, rather than automatically replacing it with 'null'. To achieve the desired behavior of serializing 'NaN' as 'null' without raising an exception, the data structure should be pre-processed to replace float 'NaN' values with 'None' (which serializes to 'null') before calling 'json.dumps'.

Suggested change
| H3 | `emergence_dashboard.py:117` | `json.dumps` with `allow_nan=True` writes invalid JSON | `allow_nan=False`; replace NaN with `null` |
| H3 | emergence_dashboard.py:117 | json.dumps with allow_nan=True writes invalid JSON | Replace NaN with None in data before json.dumps with allow_nan=False |

| H4 | `adapters.py:197-199` | Positions are contiguous prefix 1..16; docstring says "evenly-spaced" | Fix docstring (if prefix is intentional) OR implement true evenly-spaced sampling |
| H5 | `verbalizer_axes.py:249-272` | Agreement margin trivially ≥ 0; 0.5 null is vacuous | Normalize against permutation null or switch to Cohen's kappa |
| H6 | `emergence.py:115-120` | BoW/length null defaults to 0.5 on exception; weakens Tier-1 gate | Propagate exception or exclude null-failed concepts |
| H7 | `adapters.py:173-179` | 429 backoff ~9s; `Retry-After` ignored; hourly quota never clears | Honor `Retry-After`; checkpoint partial rows |
| H8 | `adapters.py:193-200` | 2 HTTP requests per `reconstruct()`; callers budget as 1 | Add request counter + pacing; fix budget comments |
| H9 | `real_nla_example.py:23` | `layer` param doesn't exist in constructor → `TypeError` on default entry | Remove `layer=` kwarg |
| H10 | `matching.py:94-98` | Lexical match is unanchored substring; "art" matches "departure" | Use word-boundary regex or strip-and-split |
| H11 | `controls.py:94-99` | Controls scored only on freq_band; token-length match ignored | Add token-length scoring term per docstring spec |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The description for H11 states that "token-length match [is] ignored" and suggests adding a token-length scoring term. However, the current implementation of build_matched_controls in nla_eval/controls.py (lines 94-97) already scores candidates using both n_tokens and char_band in addition to freq_band:

score += 1 if wp["n_tokens"] == cp["n_tokens"] else 0
score += 1 if wp["char_band"] == cp["char_band"] else 0

This indicates that the roadmap description is outdated or inaccurate relative to the existing codebase.

| H12 | `core.py:132-135` | Insertion detection never checks `ex.text`; flags source-text words as hallucinations | Include source text words in the reference set |

---

## 2. Expanded domain coverage

v0.1 evaluates on the original concept/document set (primarily deception domain).
v0.2 adds multi-domain coverage using the same NLAs, so the eval speaks to
generalization rather than domain specialization.

### 2a. Target domains for v0.2
Drawn from the ARM A balanced corpus (academically-sourced, cluster-level holdout):

| Domain | Source | Status |
|---|---|---|
| legal | Free Law Project (CourtListener) | available |
| math | AMPS / Hendrycks-MATH | available |
| medicine | MedNLI / PubMedQA | available |
| reviews | Amazon ESCI / Yelp (academic) | available |
| science | S2ORC / AllenAI | available |
| arxiv | arXiv abstracts | available |
| wiki | Wikipedia (Wikimedia dumps) | available |
| news | CC-News / RealNews | available |
| persuasion | Persuasion for Good | available |
| global_opinions | GlobalOpinionQA | available |
| deception | (v0.1 original) | existing |

Domains intentionally excluded from v0.2 expansion (data provenance issues or
insufficient academic sourcing): fineweb raw, news_rl, pku_safety, mmlu_moral.
These remain under review for v0.3.

### 2b. Concept/document set construction
- Per-domain concept set: 8 concepts minimum (the v0.1 reliability floor)
- Document pool: 10–20 documents per domain, held out from any AV training data
- Concept–document pairing: same methodology as v0.1 (canonical NLA concept list
for the tested NLA source, paired with source documents from each new domain)
- Concept survival threshold: unchanged from v0.1 (preserves backward comparability
of the threshold, not the concept set)

### 2c. Result directory structure
```
results/
v0.1/ ← v0.1 results untouched here
emergence_gemma4_deception_chunk1.json
...
v0.2/ ← all new results under this prefix
emergence_v0.2_gemma4_multidomain.json
...
```

---

## 3. Version constant + config gating

```python
# nla_eval/__init__.py
__version__ = "0.2.0.dev"
```

New domain sets live under a `BENCHMARK_VERSION` config key so a single codebase
can run either version:
```python
# nla_eval/datasets/registry.py (new file)
VERSIONS = {
"v0.1": {"domains": ["deception"], "min_concepts": 8},
"v0.2": {"domains": [...], "min_concepts": 8},
}
```

---

## 4. NLA sources (unchanged from v0.1 for now)

- `gemma-3-27b-it / kitft-l41` (Neuronpedia)
- `llama3.3-70b-it / kitft-l53` (Neuronpedia)

Adding new NLA sources (e.g., Gemma-4-E2B fine-tuned AVs from the deception
research program) is gated on: (a) the AV checkpoint being stable/published, and
(b) external review of the conditioning results. Not in v0.2 scope unless ARM
A/B/C/D produce a publishable conditioning result first.

---

## 5. Out of scope for v0.2

- Causal-fidelity (AR loop) — still P0 on the roadmap, still GPU-gated
- Human utility evals — still P1, still infra-gated
- Changing tier thresholds or the EmergenceIndex formula — would break v0.1
backward comparability; reserved for v1.0 if justified

---

## 6. Rollout gate

Per `VERSIONING.md`: branch `v0.2-dev` stays **local only** until:
- [ ] All H1–H12 harness fixes implemented and tested
- [ ] New domain concept/document sets constructed and reviewed
- [ ] At least one full eval run on a known NLA source produces a v0.2 result
- [ ] External review requested (Gemini + manual)
- [ ] Version constant set to `0.2.0` (drop `.dev`)
- [ ] Git tag `v0.2.0` pushed to remote
69 changes: 69 additions & 0 deletions docs/VERSIONING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# NLAttack versioning policy

## Guiding principle

Once a version is tagged and results are published, its benchmark content is frozen.
Changing domain coverage, NLA sources, concept sets, or scoring thresholds after
publication invalidates prior comparisons — which is worse than limited coverage.
New capabilities land in the next minor version; the prior version stays exactly
as reported.

---

## Version semantics

| Bump | When | Examples |
|---|---|---|
| **patch** (0.1.x) | Harness bug that materially affects result validity; no content change | Fix BOS-noise fallback; fix NaN JSON output; fix PYTHONHASHSEED nondeterminism |
| **minor** (0.x.0) | Additive: new domains, new NLA sources, new axes, new eval scripts | v0.2.0 — expanded domain coverage; harness fixes |
| **major** (x.0.0) | Breaking: scoring schema change, tier-threshold change, concept-set restructure that breaks backward comparability | Reserved |

**Patch rule:** a harness patch is allowed when a code bug (not a design choice)
causes published numbers to be unreliable. Patches rerun only the affected metric
on the same benchmark content and document the delta. If the delta is < 0.01 on
the published composite, the original number is noted as "confirmed within 0.01
under patch" rather than retracted.

**Minor rule:** new content always uses a new version number. The v0.1 HF card/
README gets a one-line note pointing to v0.2 for expanded coverage; it does not
change its results.

---

## Current versions

### v0.1 (released 2026-06)
- **Status: FROZEN.** Published EmergenceIndex 0.601 ("established") on
gemma-3-27b-it / kitft-l41, evaluated against the original concept/document
set. Do not change this benchmark content.
- **Known harness issues (not yet patched):**
- `emergence.py:63` — BoW buckets use `hash()` without `PYTHONHASHSEED`; Tier-1
verdict is nondeterministic across processes (does not affect v0.1 numbers if
the original eval ran in a single session with no process restart).
- `adapters.py:193-196` — schema-drift HTTP-200 silently falls back to BOS noise
instead of failing; low risk for v0.1 eval (Neuronpedia API was stable during
that run).
- `emergence_dashboard.py:117-120` — `json.dumps` with `allow_nan=True`; both
committed result files contain bare `NaN` tokens (invalid JSON for strict parsers).
Comment on lines +46 to +47

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Setting allow_nan=False in json.dumps will cause Python to raise a ValueError if any NaN values are present in the data, rather than converting them to null. To write valid JSON with null instead of NaN, the NaN values must be replaced with None in the Python dictionary/list before serialization.

- `adapters.py:197-199` — positions are a contiguous prefix 1..16, not evenly-
spaced as the docstring claims.
These are carried forward as **fixes in v0.2.0**. A v0.1 patch will be issued
only if evidence emerges that any of them affected the published composite.

### v0.2.0 (in development — LOCAL ONLY, not pushed)
See `docs/ROADMAP_v0.2.md` for scope.

---

## Rollout checklist for a new minor version

- [ ] Branch `v{N}-dev` branched from master, kept local until feature-complete
- [ ] All harness fixes from the prior version's known-issues list applied
- [ ] New content (domains, NLA sources, concept sets) added under a new config
key so both versions can run from the same codebase
- [ ] `results/` subdirectory named `v{N}/` — never overwrite prior version results
- [ ] Version constant bumped in `nla_eval/__init__.py`
- [ ] HF card updated: new version section added; prior version section unchanged
- [ ] CITATION.cff `version` field updated
- [ ] Git tag `v{N}.{M}.{P}` pushed after external review (not before)
- [ ] Prior version README note: "See v0.2 for expanded domain coverage."
1 change: 1 addition & 0 deletions nla_eval/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@
Plug in any NLA via the `NLA` adapter contract (one method: reconstruct), run a
tagged dataset through `core.run`, then read off the 20 tests as group-bys.
"""
__version__ = "0.2.0.dev" # v0.1 benchmark content frozen; see docs/VERSIONING.md
from .adapters import NLA, MockNLA, CallableNLA, NeuronpediaNLA, KitftNLA
from .matching import Matcher, EnsembleMatcher
from .core import Example, run, RunResult
Expand Down