-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy path.env.example
More file actions
1857 lines (1605 loc) · 92.2 KB
/
Copy path.env.example
File metadata and controls
1857 lines (1605 loc) · 92.2 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
# Example environment for ecaa-workflow.
# Copy to `.env` (gitignored) and fill in the values you need.
#
# Only the server loads `.env` automatically (via dotenvy in
# crates/server/src/main.rs). The harness and CLI read env vars
# from the live shell — `source .env` before invoking them, or rely
# on `make` targets that pass the vars through.
#
# =============================================================================
# Local production baseline (recommended active settings)
# =============================================================================
# A local production deployment should set only the following active values and
# leave everything else at its code default unless a specific operator need
# applies. The defaults below are durable, loopback-only, and avoid live test,
# eval, debug, remote-compute, and external-validator paths.
#
# ECAA_ANTHROPIC_API_KEY=sk-ant-... # optional for chat-side LLM; unset uses MockLlmBackend
# # (legacy ANTHROPIC_API_KEY accepted with deprecation warning)
# ECAA_EXECUTOR_MODE=local # or `aws` / `slurm` (see below)
# ECAA_HARNESS_CONCURRENCY=1 # or `auto` on a multi-core host
# ECAA_SESSION_TOKEN_BUDGET=500000 # per-session uncached-input ceiling
#
# AWS or SLURM deployments additionally need the full required block from the
# matching section below. Dev/test knobs (`ECAA_LIVE_API`, `PLAYWRIGHT_LIVE`,
# `KEEP_PACKAGE`, `ECAA_DEBUG_TOKEN_BURN`, `ECAA_MOCK_DISCOVERY_BLOCK`) should
# NOT be set in production.
# =============================================================================
# LLM backend
# =============================================================================
# Required for the live Anthropic path (chat server, chat-llm CLI, nightly
# scorer, live Playwright specs). When unset the server falls back to
# MockLlmBackend and chat-llm / live tests SKIP.
#
# NOTE: This key is for the CHAT LLM only (the UX-shim conversation layer).
# The per-task Claude Code EXECUTOR agent defaults to SUBSCRIPTION billing
# at ~/.claude/.credentials.json (Max/Pro plan). The SWFC-prefixed name
# prevents accidental collisions with Claude Code CLI's ANTHROPIC_API_KEY
# scan. Legacy `ANTHROPIC_API_KEY` is still accepted (with a one-time
# stderr deprecation warning) to keep existing .env files working.
# See ECAA_AGENT_BILLING below.
# ECAA_ANTHROPIC_API_KEY=
# Override the Anthropic API base URL (proxy, staging). Default: https://api.anthropic.com
# ANTHROPIC_BASE_URL=https://api.anthropic.com
# Agent billing mode. Controls whether scripts/agent-claude.sh invokes
# `claude-code` against the subscription (default) or the API.
# subscription — unset ECAA_ANTHROPIC_API_KEY for the agent; use
# ~/.claude/.credentials.json. Per-task cost: $0
# (subscription quotas apply). Default.
# api — forward ECAA_ANTHROPIC_API_KEY into the agent env; each
# task bills the API at token × rate. Use only for CI
# or hosts without a subscription.
ECAA_AGENT_BILLING=subscription
# Executor LLM backend. The harness invokes scripts/agent.sh (a dispatcher);
# this selects which wrapper it execs. The harness↔agent contract is
# file-based (the wrapper writes result.json/state.patch.json/.heartbeat), so
# the backend is invisible to the harness binary.
# claude (default) → scripts/agent-claude.sh (Claude Code; production path)
# codex → scripts/agent-codex.sh (Codex CLI, `codex exec --yolo`)
# Codex CLI delivery mirrors the Claude path: agent-codex.sh runs a host-side
# `npm install @openai/codex` into a cache dir and bind-mounts its node_modules
# into the container (the codex native binary is static musl → runs in bio-min;
# the image already has the node runtime). Codex auth: ECAA_OPENAI_API_KEY is
# the rotation-free headless path; absent that, a ChatGPT login at
# ECAA_CODEX_AUTH_DIR (default ~/.codex) is copied (writable, refreshable) into
# the per-task agent HOME. NOTE: a ChatGPT-login auth.json carries no usable
# OPENAI_API_KEY, so the key path needs a real platform API key.
# ECAA_AGENT_BACKEND=claude
# ECAA_AGENT_CODEX_EXPERIMENTAL=0 # 1 → opt in to the EXPERIMENTAL codex backend;
# # agent-codex.sh refuses to run unless this is 1
# # (no per-task images/GPU/cost-telemetry parity yet)
# ECAA_OPENAI_API_KEY= # platform API key (not the ChatGPT-login token)
# ECAA_CODEX_AUTH_DIR=~/.codex # ChatGPT-login dir; default already ~/.codex
# ECAA_AGENT_CODEX_MODEL=gpt-5.5 # codex model; ECAA_AGENT_MODEL_OVERRIDE wins
# ECAA_AGENT_CODEX_VERSION=latest # npm @openai/codex version to install
# ECAA_AGENT_CODEX_DISABLE=0 # 1 → use the image's bundled codex
# ECAA_AGENT_CODEX_FORCE_REINSTALL=0 # 1 → reinstall even if cached
# ECAA_CODEX_INSTALL_DIR= # explicit cache dir (else session/agent cache)
# ECAA_AGENT_HOME_DIR= # override the per-task agent HOME (both wrappers).
# # Default is an OUT-OF-PACKAGE cache dir so agent
# # credentials never land inside the emitted package
# # (artifact-served + git-committed). Overriding it
# # back into $PACKAGE re-opens that exposure.
# ECAA_TASK_BUDGET_USD= # per-task soft dollar budget the codex wrapper passes
# # into the container (the claude path uses the CLI's
# # native --max-budget-usd instead); 0/unset disables.
# Per-task model tiering. With =1 (default), agent routes by task class:
# discover_* → Opus 4.8 (analytical scoring; depth pays off)
# validate_* → Sonnet 4.6 (stable codegen shape; 5× cheaper than Opus)
# data_acq → Sonnet 4.6
# data_import → Sonnet 4.6
# everything else → Sonnet 4.6
# Quality safety net: validate_* still runs after every analytical task
# with explicit integrity checks; Sonnet writes stable validation
# scripts reliably. Subscription billing benefits too — Anthropic's
# quota meters token consumption regardless of mode. Set =0 to revert
# to all-Opus when maximum quality is required regardless of cost.
# ECAA_AGENT_MODEL_TIER=0
# Force every agent task to a single Claude model id, bypassing per-task
# model tiering. Intended for eval fairness or controlled operator
# calibration; unset in production so task-class tiering applies.
# ECAA_AGENT_MODEL_OVERRIDE=claude-sonnet-4-6
# Hard per-task cost ceiling (USD). claude CLI's --max-budget-usd
# terminates the agent session when running cost crosses the limit;
# this replaces the prior post-hoc enforce_turn_budget_limit (which
# only blocked AFTER the agent had already spent the tokens).
# Set above the cross-MODALITY p99 of each class's real cost distribution
# (on the model that runs the class) — a cap is a ceiling, and cost varies
# by modality, so calibrate from a variety of runs, not one workflow.
# Defaults below come from 1243 task runs across 12 modalities. discover_*
# (Opus) is the expensive class — bulk_rnaseq/variant_calling discover cost
# ~$2-2.7, so a low discover cap false-blocks ~30% of them.
#
# Per-class overrides (defaults in scripts/agent-claude.sh):
# ECAA_AGENT_BUDGET_USD_VALIDATE=1.25
# ECAA_AGENT_BUDGET_USD_DISCOVER=3.00
# ECAA_AGENT_BUDGET_USD_DATA_ACQ=2.00
# ECAA_AGENT_BUDGET_USD_ANALYTICAL=3.00 # also governs reporting/final_reporting
#
# Global override beats per-class. Set to 0 to disable the cap entirely.
# ECAA_AGENT_BUDGET_USD=3.00
# Verbose tracing inside scripts/agent-claude.sh (dispatch context, mount
# inspection, credential refresh). Default: 0.
# ECAA_AGENT_DEBUG=1
# =============================================================================
# Chat server
# =============================================================================
# Bind interface. Default 127.0.0.1 (loopback). 0.0.0.0 exposes on LAN;
# the server REFUSES TO BOOT on a non-loopback bind unless either
# ECAA_SERVER_AUTH_TOKEN is set or ECAA_SERVER_LAN_NO_AUTH=1 is the
# explicit opt-in.
# ECAA_BIND_ADDR=127.0.0.1
# HTTP listen port. Default 3000. Must match ECAA_SERVER_URL's port and
# the UI proxy's VITE_API_PORT — the `dev_port_drift` test guards this.
# ECAA_PORT=3000
# Bearer token required for every request when the server binds anything
# other than 127.0.0.1, and accepted as an opt-in shared secret on
# loopback. Generate with `openssl rand -hex 32` and pass via
# `Authorization: Bearer <token>` from the UI / curl. Unset on loopback
# = no auth required (single-user dev). Required for non-loopback bind.
# ECAA_SERVER_AUTH_TOKEN=
# Explicit opt-in to bind LAN-exposed (non-loopback) WITHOUT a token.
# Without this, the server fails closed on a non-loopback bind that
# lacks ECAA_SERVER_AUTH_TOKEN. Do not use in production — only for
# isolated dev VLANs.
# ECAA_SERVER_LAN_NO_AUTH=1
# Per-session harness self-token. NOT operator-set: the server mints a
# random token at /start_execution, stores only its SHA-256 on the
# ExecutionHandle, and injects the raw value into the harness child via
# this env var. The harness sends it back as `X-Harness-Token`, so its
# requests are principal'd as a session-scoped HarnessAgent rather than
# the global admin Owner. Documented here only to satisfy the env-var
# doc gate; setting it manually has no useful effect.
# ECAA_HARNESS_TOKEN=
# Allowed CORS Origins (comma-separated). Default covers the dev UI on
# :5173 + :3000 in both 127.0.0.1 and localhost forms. Override to the
# actual UI host(s) for production deployments.
# ECAA_CORS_ORIGINS=http://localhost:5173,http://127.0.0.1:5173,http://localhost:3000,http://127.0.0.1:3000
# Disable per-session owner authorization (debug only). With =1, any
# authenticated request can read/write any session. Used by the
# crates/server/tests/owner_authz.rs bypass branch. Never in production.
# ECAA_OWNER_AUTHZ_DISABLE=1
# Bypass the per-server flock (RC-22 multi-server tests). Default off;
# setting =1 lets a second server attach to a session already locked by
# another instance. Tests only.
# ECAA_SERVER_DEBUG_ALLOW_MULTI_PROCESS=1
# `offline` forces the server onto MockLlmBackend even if ECAA_ANTHROPIC_API_KEY is
# set. Kill switch for UI work when the LLM is unavailable.
# ECAA_CHAT_MODE=offline
# Directory where chat sessions are persisted. Default: ~/.ecaa-workflow/sessions.
# Keep production sessions on durable storage; /tmp is for disposable tests only.
ECAA_CHAT_SESSIONS_DIR=$HOME/.ecaa-workflow/sessions
# Config root used by the chat server (taxonomies, modality keywords, policies).
# Default: ./config (resolved relative to CWD).
# ECAA_CONFIG_DIR=./config
# Override the search root when the chat tools look up emitted packages by name.
# Default: ~/.ecaa-workflow/packages
ECAA_PACKAGE_ROOT=$HOME/.ecaa-workflow/packages
# Override the path to config/model-policy.yaml for ad-hoc experimentation.
# ECAA_MODEL_POLICY_PATH=/path/to/model-policy.yaml
# Override the confidence threshold at which the model policy escalates
# Sonnet → Opus. Default sourced from model-policy.yaml (0.3). Must be
# a float in [0.0, 1.0]; out-of-range or non-numeric values are ignored
# with a stderr warning.
# ECAA_MODEL_ROUTING_CONFIDENCE_THRESHOLD=0.3
# Session-store load mode. permissive (default) skips a corrupted file
# with a warning; strict fails the whole load. Test sites bias to strict.
# ECAA_SESSION_LOAD_MODE=strict
# Cadence of the session-store TTL prune sweep (seconds). Default 3600
# (1 hour). The prune removes sessions inactive longer than 30 days.
# ECAA_TTL_PRUNE_INTERVAL_SECS=3600
# Per-request body size cap (KiB). Default 256. Bumped automatically by
# the streaming upload code path (which has its own ECAA_UPLOAD_*
# bounds), so this cap is only the JSON-RPC ceiling.
# ECAA_REQUEST_BODY_LIMIT_KB=256
# Per-request wall-clock timeout (seconds). Default 60. SSE streams
# bypass this guard.
# ECAA_REQUEST_TIMEOUT_SECS=60
# Idempotency-Key cache TTL (seconds). Default 3600 (1 hour). Replays
# inside this window with the same body short-circuit to the cached
# response — see crates/server/src/chat_routes/_idempotency.rs.
# ECAA_IDEMPOTENCY_TTL_SECS=3600
# Per-endpoint LLM rate limits (requests / minute, per-IP via
# tower-governor). All have safe production defaults; bump in dev when
# the UI's polling endpoints saturate the global limit and need
# isolation. See crates/server/src/chat_routes/_rate_limits.rs.
# ECAA_LLM_RATE_LIMIT_TURN POST /turn (default 30)
# ECAA_LLM_RATE_LIMIT_SCORE scoring side-call (default 6)
# ECAA_LLM_RATE_LIMIT_EXPLAIN explain side-call (default 30)
# ECAA_LLM_RATE_LIMIT_SUMMARY summary side-call (default 6)
# ECAA_LLM_RATE_LIMIT_REMEDIATION remediation-proposer (default 6)
# ECAA_LLM_RATE_LIMIT_START_EXEC POST /start-execution (default 12)
# ECAA_LLM_RATE_LIMIT_BRANCH POST /branch (default 6)
# ECAA_LLM_RATE_LIMIT_SME_EDIT SME task-param/bound edit (default 30)
# ECAA_LLM_RATE_LIMIT_TURN=30
# ECAA_LLM_RATE_LIMIT_SCORE=6
# ECAA_LLM_RATE_LIMIT_EXPLAIN=30
# ECAA_LLM_RATE_LIMIT_SUMMARY=6
# ECAA_LLM_RATE_LIMIT_REMEDIATION=6
# ECAA_LLM_RATE_LIMIT_START_EXEC=12
# ECAA_LLM_RATE_LIMIT_BRANCH=6
# ECAA_LLM_RATE_LIMIT_SME_EDIT=30
# Read-only shared-URL middleware. =1 honors session-share tokens for
# unauthenticated browsers. Default 0; without it the middleware is a
# no-op (every request passes through).
# ECAA_SHARED_URLS_ENABLED=1
# Disposition auto-apply (TEST ONLY). =1 auto-applies disposition decisions
# the agent emits, skipping SME confirmation. Do not set in production —
# the disposition unit tests intentionally clear it.
# ECAA_AUTO_APPLY_DISPOSITIONS=1
# Input listing + upload bounds (chat-routes inputs/upload routes).
# Tune when SMEs work with large cohort directories or upload-driven
# pipelines saturate the server's disk reserve.
# - ECAA_INPUT_ROOTS=<colon-separated>: allowlist of filesystem roots
# SMEs may point local_path inputs at. Default user-home only.
# - ECAA_INPUT_MAX_FILES=<n>: max files listed (default 50000).
# - ECAA_INPUT_MAX_FILE_BYTES=<n>: per-file byte cap (default 50 GB).
# - ECAA_INPUT_MAX_TOTAL_BYTES=<n>: total bytes across listed files
# (default 250 GB).
# - ECAA_UPLOAD_ROOT=<dir>: root for uploaded session files
# (default ~/.ecaa-workflow/uploads).
# - ECAA_UPLOAD_MAX_BYTES=<n>: per-upload byte cap.
# - ECAA_UPLOAD_DISK_RESERVE_GB=<n>: bail out of new chunk writes when
# free space on the upload root's filesystem falls below this (default 50).
# ECAA_INPUT_ROOTS=$HOME
# ECAA_INPUT_MAX_FILES=50000
# ECAA_INPUT_MAX_FILE_BYTES=53687091200
# ECAA_INPUT_MAX_TOTAL_BYTES=268435456000
# ECAA_UPLOAD_ROOT=~/.ecaa-workflow/uploads
# ECAA_UPLOAD_MAX_BYTES=53687091200
# ECAA_UPLOAD_DISK_RESERVE_GB=50
# Package import (upload & explore) bounds — zip/tar-bomb defense for the
# server package-import endpoint.
# Max uploaded package archive size in bytes (default 2 GiB).
ECAA_MAX_IMPORT_BYTES=2147483648
# Max entries in an uploaded package archive (default 200000).
ECAA_MAX_IMPORT_ENTRIES=200000
# Max total decompressed size of an uploaded package archive in bytes
# (default 8 GiB). Decompression-bomb defense: the compressed-upload cap
# above can't bound how large a tiny archive expands to, so the extractor
# caps the running decompressed total too.
ECAA_MAX_IMPORT_EXTRACTED_BYTES=8589934592
# Auto-register filesystem path hints extracted from intake prose
# (chat-side `append_intake_prose`). When `=1`, any path-shaped token
# in the SME's prose that resolves under ECAA_INPUT_ROOTS is promoted
# directly onto `session.inputs` — no follow-up confirmation prompt.
# Useful in non-interactive fixture runs where no SME is present to
# approve the registration. Default off (interactive flow surfaces the
# hint via get_session_state + the UI offers to register).
# ECAA_AUTO_REGISTER_PROSE_PATHS=1
# Anthropic client + token-burn knobs (chat side only; the per-task
# agent runs through scripts/agent-claude.sh's own credential path).
# - ECAA_ANTHROPIC_TIMEOUT_SECS=<n>: HTTP timeout for the Messages API
# (default 180; was 120 pre plan §S2.7). Long generations on
# Sonnet/Opus + tool loops can exceed 60–90s.
# - ECAA_ANTHROPIC_RPM=<n>: requests-per-minute cap on outbound calls.
# Default unset (no cap). Set when the deployment hits a rate-limit.
# - ECAA_DISABLE_CONTEXT_EDITING=1: disables `context-management-2025-06-27`
# (auto-clears stale tool_result blocks after 8 iterations; default on).
# - ECAA_SLIM_TAXONOMY=1: strips per-stage descriptions from the
# taxonomy prompt to shrink uncached input.
# - ECAA_SESSION_TOKEN_BUDGET=<n>: soft-blocks the tool loop at this
# many uncached input tokens (default 500000; cache-read doesn't
# bill, so the budget tracks fresh content only).
# - ECAA_BUDGET_HARD_STOP=1: with ECAA_SESSION_TOKEN_BUDGET set, the
# tool loop refuses to start a new turn once the soft-block ceiling
# is crossed (default warn-only). CI fixtures use this to gate spend.
# - ECAA_DEFAULT_SESSION_BUDGET_USD=<n>: server-side default
# session-level cost cap (USD). Per-session overrides via the
# Settings UI; this is the fall-through default.
# - ECAA_ALLOW_1H_CACHE=1: permits ttl: "1h" on cache_control entries.
# Default rejected (5m only).
# - ECAA_DEBUG_TOKEN_BURN=1: streams per-turn token/cache/cost to
# stderr, warns at cache-hit <30% after ≥3 turns. Debug only.
# - ECAA_DEBUG=1: master debug gate. ECAA_DUMP_ANTHROPIC_PAYLOAD is
# ignored unless this is set (refuses-by-default so a forgotten dump
# path doesn't leak prompts into operator log directories).
# - ECAA_DUMP_ANTHROPIC_PAYLOAD=<file>: writes JSON of every outbound
# Messages API request (mode 0600, appended). Use per-session paths
# like /tmp/ecaa-payload-$SESSION.jsonl. Gated on ECAA_DEBUG=1.
# ECAA_ANTHROPIC_TIMEOUT_SECS=180
# ECAA_ANTHROPIC_RPM=50
# ECAA_DISABLE_CONTEXT_EDITING=1
# ECAA_SLIM_TAXONOMY=1
ECAA_SESSION_TOKEN_BUDGET=500000
# ECAA_BUDGET_HARD_STOP=1
# ECAA_DEFAULT_SESSION_BUDGET_USD=50
# ECAA_ALLOW_1H_CACHE=1
# ECAA_DEBUG_TOKEN_BURN=1
# ECAA_DEBUG=1
# ECAA_DUMP_ANTHROPIC_PAYLOAD=/tmp/ecaa-anthropic-payload.jsonl
# Git-backed provenance.
# - ECAA_GIT_ENABLED=0 hard-disables the feature regardless of the
# Settings UI checkbox. Any other value (or absent) is config-driven
# (default off in ~/.ecaa-workflow/git-config.json).
# - ECAA_GIT_CONFIG_PATH=<path>: override config-file path (default
# ~/.ecaa-workflow/git-config.json).
# - ECAA_GIT_REPO_ROOT=<absolute-path>: operator escape hatch.
# GitConfig::validate requires repo_path to canonicalize under $HOME;
# set this to an additional absolute root that's also acceptable
# (e.g., a shared /var/scripps/packages mount).
ECAA_GIT_ENABLED=1
# ECAA_GIT_CONFIG_PATH=~/.ecaa-workflow/git-config.json
# ECAA_GIT_REPO_ROOT=/var/scripps/packages
# ECAA_COMPOSER — planner selection. `semantic` and `proof-carrying` both
# select the v4 proof-carrying planner (emits runtime/proofs.jsonl +
# runtime/assumptions.jsonl; persisted onto Session::composer_version). Legacy
# values (`legacy`/`archetypes`/`backward-chain`) warn and route to v4. Default
# is the v4 planner when unset.
# ECAA_COMPOSER=proof-carrying
# ECAA_COMPOSE_STRICT — when set truthy, the v4 composer runs in
# RiskMode::Production: every composed-DAG edge that is not a typed data
# flow or adapter-mediated rejects, INCLUDING archetype-declared
# ordering-only edges. Default off (Draft) — clinical-tier emission that
# wants ordering exemptions re-justified sets this on.
ECAA_COMPOSE_STRICT=1
# ECAA_COMPOSE_INTERPRETATION — when set truthy, the v4 composer injects a
# grounded biological_interpretation atom (biology + method-justification,
# optionally literature-contextualized) before every reporting terminal, and a
# validate_interpretation companion. Feature-flagged (A/B-gated); default OFF,
# but opted IN here so the composer emits grounded interpretation. Unset (or =0)
# to revert to unchanged emitted DAGs.
ECAA_COMPOSE_INTERPRETATION=1
# ECAA_INTAKE_RESOLUTION — how intake resolves when NO specific analysis atom
# binds to the SME's question (the silent generic-fallback failure mode where a
# DAG finalises with only the descriptive scaffold and zero bound analysis).
# auto-author (DEFAULT) — splice an `agent_generated_analysis` node so the
# run ATTEMPTS the requested analysis instead of emitting a
# vacuous descriptive-only plan. Aliases: author, (unset/empty).
# sme — block at emit and wait for the human SME to author/confirm.
# strict-block — refuse emit and record an unbindable-analysis failure.
# Aliases: strict, block.
# Only acts when the question actually requests an analysis (an analysis verb is
# present); a genuinely descriptive ask is left on the generic scaffold.
# ECAA_INTAKE_RESOLUTION=auto-author
# ECAA_EXTERNAL_CURATED_DIRS — colon-separated operator-declared curated
# external-registry snapshot dirs (F3 federation). Dirs in this list
# resolve at RegistryTier::Curated (may reach StaticChecked after
# validate_for_executable); everything else is RegistryTier::Community
# (capped at Unverified). Empty by default. Imports never reach production
# execution until promoted regardless of tier (the execution gate is the
# backstop).
ECAA_EXTERNAL_CURATED_DIRS=
# =============================================================================
# Compute backend (harness)
# =============================================================================
# Which Executor to use.
# local — runs the agent as a subprocess on this host (default)
# aws — provisions EC2 via the aws CLI (see AWS section below)
# slurm — submits sbatch jobs over SSH to a SLURM cluster (see SLURM section below)
# `mock` exists as a test-only constructor; not dispatched from the factory.
ECAA_EXECUTOR_MODE=local
# =============================================================================
# Literature retrieval (literature_context tool + literature_* DAG tasks)
# =============================================================================
# Source-scope tier for evidence retrieval. Three values:
# pmc_oa — PubMed Central Open Access subset (default)
# pmc_oa_plus_abstracts — also pulls PubMed/E-utilities abstracts
# all_sources_local_only — phase-2 stub (institutional access)
# Invalid values fall back to pmc_oa with a stderr warning.
# ECAA_LIT_SOURCE_SCOPE=pmc_oa_plus_abstracts
# Method-source authority for literature-grounded discovery. Values:
# frozen — rank only the already-curated candidate pool
# bounded — default; live discovery bounded by retrieval routes
# open_ended — reserved wider mode; currently behaves like bounded
# ECAA_METHOD_SOURCE_AUTHORITY=bounded
# NCBI E-utilities API key. Unset = 3 req/s. Set = 10 req/s. Never logged.
# ECAA_LIT_NCBI_API_KEY=
# Per-task literature evidence size cap (MB). Default 200. On hit,
# result.json carries `truncated_at_storage_cap: true`. Soft-truncate,
# not a blocker.
# ECAA_LIT_EVIDENCE_MAX_MB=200
# Phase-2 stub for `all_sources_local_only`. No-op when set against any
# other scope. Not for production.
# ECAA_LIT_INSTITUTIONAL_ACCESS=1
# Disable the literature_context tool entirely. Used by external-eval
# adapters (Biomni, BioMedAgent) to force parity comparisons without
# the literature shim. Default off.
# ECAA_LITERATURE_CONTEXT_DISABLED=1
# Playwright literature.spec.ts live-API gate. Default off (the spec
# SKIPS without it).
# ECAA_LIT_LIVE_API=1
# =============================================================================
# AWS provisioning — required when ECAA_EXECUTOR_MODE=aws
# =============================================================================
# The required AWS variables are documented in this section. Every var below is
# no-op when ECAA_EXECUTOR_MODE=local.
# EC2 region for provisioning.
# ECAA_AWS_REGION=us-west-2
# AMI the harness launches. Produce with the Packer template in scripts/packer/.
# ECAA_AWS_AMI_ID=ami-0123456789abcdef0
# Security group attached to provisioned instances. Must allow outbound 443 for
# SSM and S3.
# ECAA_AWS_SECURITY_GROUP=sg-0123456789abcdef0
# IAM instance profile — required for SSM RunCommand + S3 access.
# ECAA_AWS_INSTANCE_PROFILE=ecaa-workflow-agent
# VPC subnet(s). Set either a single subnet OR a comma-separated list for
# multi-AZ failover (multi_az_policy rotates on InsufficientInstanceCapacity).
# ECAA_AWS_SUBNET_ID=subnet-0123456789abcdef0
# ECAA_AWS_SUBNET_IDS=subnet-aaa,subnet-bbb,subnet-ccc
# Optional EC2 key pair for SSH break-glass debugging. SSM is the primary
# control channel.
# ECAA_AWS_KEY_PAIR=scripps-ops
# S3 bucket + key prefix used to stage package inputs and collect outputs.
# Defaults: ecaa-workflow / ecaa-workflow/
# ECAA_AWS_S3_BUCKET=ecaa-workflow
# ECAA_AWS_S3_PREFIX=ecaa-workflow/
# Short SHA of the harness workspace. Tagged onto every provisioned instance
# so scan_orphans can correlate. Default: "unknown".
# ECAA_WORKSPACE_SHA=$(git rev-parse --short HEAD)
# =============================================================================
# AWS policy knobs
# =============================================================================
# true = request spot capacity with CapacityRebalance=true. Default: false.
# ECAA_AWS_SPOT=true
# REQUIRED when ECAA_EXECUTOR_MODE=aws. Hard USD cap — provision aborts if the
# estimated run cost exceeds the ceiling. Phase 9 (Task 9.1, closes F-EXEC-C-02):
# the harness now fails closed without an explicit positive ceiling. Operators
# that don't want a ceiling at all should run the non-AWS executors (local /
# slurm) which use NoopCostModel.
# ECAA_AWS_COST_CEILING_USD=10.00
# Phase 9 (Task 9.2) — optional comma-separated allowlist of EC2 instance
# types the harness is permitted to launch. Unset = any instance type the
# sizing layer chooses. Recommended starting set covers the workloads the
# default sizing tables emit; tighten further for prod operators.
# ECAA_AWS_INSTANCE_TYPE_ALLOWLIST=t3.medium,m6i.large,m6i.xlarge,m6i.2xlarge,r6i.large,r6i.xlarge,c6i.large,c6i.xlarge,g5.xlarge,g5.2xlarge
# What to do when a task exceeds its sizing high-water mark.
# block — halt the task
# resize — upgrade to the next shape that fits (default)
# continue — warn and run anyway
# ECAA_AWS_HIGH_WATER_POLICY=resize
# Orphan instance policy.
# warn — log orphaned i-XXXX without terminating (default)
# reap — actively terminate orphans
# Any other value is treated as warn. CI lanes should set =reap so a
# leaked instance from a prior flaky run is terminated automatically.
# ECAA_AWS_ORPHAN_POLICY=warn
# Per-task SSM command timeout (seconds). Default: 3600. Per-stage overrides
# come from policies/compute-resource-policy.json inside each emitted package.
# ECAA_AWS_SSM_TIMEOUT_SECS=3600
# Verified-orphan sweep deadline (seconds). The harness times out the
# secondary describe-instances confirmation step after this long.
# Default: 300.
# ECAA_AWS_ORPHAN_VERIFY_TIMEOUT_SECS=300
# Regional pricing multiplier applied to the bundled instance-type
# pricing table. Default 1.0 (us-west-2 reference). Validated as a
# finite float in [0.5, 5.0] at config load.
# ECAA_AWS_PRICING_REGION_MULT=1.10
# Override the bundled pricing table. Accepts EITHER inline JSON OR a
# path to a JSON file. Schema: { "instance-type": <USD/hr>, ... }.
# Unspecified types continue to use bundled × ECAA_AWS_PRICING_REGION_MULT.
# ECAA_AWS_PRICING_OVERRIDES_JSON=/etc/scripps/aws-pricing.json
# ECAA_AWS_PRICING_OVERRIDES_JSON={"r6i.8xlarge":2.10}
# Spot bid price as a fraction of on-demand. Default 0.30. Higher
# values reduce spot-loss probability at higher cost.
# ECAA_AWS_SPOT_DISCOUNT_FRACTION=0.30
# Cumulative spend cap across the entire harness run (USD). Default
# 100. Differs from ECAA_AWS_COST_CEILING_USD which gates a single
# dispatch. On breach the harness halts and surfaces
# BlockerKind::CostCeilingExceeded.
# ECAA_AWS_RUN_TOTAL_CEILING_USD=100.00
# Security group for tasks whose atom carries `safety.network: none`.
# Required when the DAG contains any such task — the harness refuses
# dispatch otherwise. The SG should drop all outbound traffic.
# ECAA_AWS_RESTRICTED_SG_ID=sg-0fedcba9876543210
# =============================================================================
# Agent-side (used by scripts/agent-claude-aws.sh on the provisioned host)
# =============================================================================
# AWS credential profile for the CloudWatch MCP server. Default: default
# ECAA_AWS_PROFILE=default
# Populated by the AwsExecutor at dispatch time so the agent can tag
# CloudWatch queries with the right InstanceId dimension. Do not set by hand.
# ECAA_AWS_INSTANCE_ID=i-0123456789abcdef0
# Test-only handles into the AWS executor's command-cache. Not for
# production use; the integration tests set these to pin behavior.
# ECAA_AWS_ENV_LOCK=1
# ECAA_AWS_COMMAND_ID=cmd-XXXX
# Toggle MCP-server blocks in the agent's .mcp.json. All default to 0.
# Set =1 to enable the corresponding MCP server for an AWS-side agent.
# ECAA_MCP_AWS_AGENT_REGISTRY=1
# ECAA_MCP_BIO=1
# ECAA_MCP_GITHUB=1
# =============================================================================
# Harness scheduler
# =============================================================================
# Parallel-dispatch budget for the harness scheduler. `1` (default) is
# serial behavior. `auto` resolves to `nproc / max(tool_thread_curves)`
# so each task saturates its own thread budget without oversubscription.
# Integers > 1 cap the pool at that size. See crates/harness/src/scheduler.rs.
ECAA_HARNESS_CONCURRENCY=1
# Validation-lane parallelism (numeric 0/1). `1` = primary executor
# advances processing/analysis tasks while a secondary executor
# independently runs any ready validate_* task. `0` (or any other
# value) keeps validators serial. Overrides ECAA_HARNESS_CONCURRENCY
# when set to `1`. True parallel execution requires
# ECAA_EXECUTOR_MODE=local; aws/slurm keep the lane picker correctness
# but serialise execution through one backend handle (avoids
# double-provisioning a remote instance / submitting two batch jobs
# for one logical lane).
# ECAA_HARNESS_VALIDATION_LANE=1
# Harness loop tuning (advanced — leave at defaults unless diagnosing).
# - ECAA_HARNESS_AMEND_CANCEL=1 kills an in-flight task whose stage was
# just amended, instead of letting it finish (default off; finishing
# the prior attempt is usually cheaper than re-running from scratch).
# - ECAA_HARNESS_SETTLE_SECS controls the quiet period the harness
# waits between iterations after a task completes (default 1s; bump
# on slow filesystems where the agent's output write isn't visible
# to the harness scanner immediately).
# - ECAA_HARNESS_BATCH_WINDOW_SECS — chat-side debounce window before
# the HarnessBatcher flushes accumulated harness-progress events as
# one synthetic assistant turn (default 10s; out-of-range / 0 /
# non-numeric values fall back to default; values >600 are rejected
# so blockers stay visible).
# - ECAA_TASK_HEARTBEAT_STALL_SECS — heartbeat stall threshold
# (default 300 = 5 min, per DEFAULT_TASK_HEARTBEAT_STALL_SECS in
# crates/core/src/config.rs:57; set 0 to disable). On breach the
# harness flips the Running task to Blocked { HeartbeatStalled }.
# - ECAA_HEARTBEAT_LIVENESS_SECS — orphan-by-crash recovery liveness
# window (default 60s, clamped [0, 600]).
# - ECAA_HARNESS_FINALIZE_PROBE_MIN_INTERVAL_SECS / _TIMEOUT_SECS —
# finalize-probe cadence + per-probe deadline for silent-completion
# detection (defaults 30s / 10s).
# - ECAA_HARNESS_MAX_NOPROGRESS_DISPATCHES — consecutive dispatches that
# produce no progress before the crash-loop guard force-blocks the task
# (default 8; set 0 to disable).
# - ECAA_HARNESS_VALIDATION_RECOVERY — DEFAULT OFF. When truthy
# (1/true/yes/on), a task that a `required` validation-contract assertion
# re-blocked is re-dispatched to the SAME agent a bounded number of times
# after the harness writes a METHOD-NEUTRAL domain-correctness signal
# (the failed assertion id + the design's operator-authored bound vs the
# number recomputed from the agent's OWN result.json — never a tool, flag,
# or threshold value) into runtime/inputs/<task>/domain-correctness-signal.json.
# Off by default so the production / SME path keeps its human checkpoint
# (a blocked task stays blocked for the SME). Intended for the unattended
# operator-run eval arm; the per-task signal file records the recovery
# budget so the scorecard meta can disclose that the bare arm has no
# retry loop. The budget is durable on disk so it stays bounded across
# the server's harness auto-relaunch.
# - ECAA_HARNESS_VALIDATION_RECOVERY_MAX — per-task recovery attempt budget,
# clamped to [0, 2] (default 2). 0 disables recovery even when the flag
# above is on.
# - ECAA_HARNESS_CONTRACT_ADVISORY — DEFAULT OFF. When truthy
# (1/true/yes/on), ALL domain-correctness gates become NON-BLOCKING
# advisory diagnostics: the task stays in its completed state so the DAG
# proceeds and no validation-recovery re-dispatch fires. Covers three
# gates — (1) the harness `required` validation-contract assertions and
# (2) the harness Phase-13 post-completion validator bundle (both append
# one JSON line per failure — task_id, assertion_id, severity, reason —
# to a per-package runtime/validation-warnings.jsonl sidecar) and (3) the
# server-side claim_coverage recall-gap re-verify gate (the gap is already
# persisted into the signed verdict sink + audit-proof report, so the
# Blocked/ValidationFailed transition is simply suppressed). Takes
# PRECEDENCE over ECAA_HARNESS_VALIDATION_RECOVERY — when both are set,
# advisory wins (no block, no re-dispatch). Off by default so the
# production / SME path keeps its human checkpoint (a failed required
# assertion or recall gap blocks the task for the SME). Intended for a
# fair single-shot eval arm that runs the strict domain-correctness checks
# as diagnostics without dead-stalling tasks or triggering recovery retries.
# ECAA_HARNESS_AMEND_CANCEL=1
# ECAA_HARNESS_SETTLE_SECS=1
# ECAA_HARNESS_BATCH_WINDOW_SECS=10
# ECAA_TASK_HEARTBEAT_STALL_SECS=300
# ECAA_HEARTBEAT_LIVENESS_SECS=60
# ECAA_HARNESS_FINALIZE_PROBE_MIN_INTERVAL_SECS=30
# ECAA_HARNESS_FINALIZE_PROBE_TIMEOUT_SECS=10
# ECAA_HARNESS_MAX_NOPROGRESS_DISPATCHES=8
# ECAA_HARNESS_VALIDATION_RECOVERY=0
# ECAA_HARNESS_VALIDATION_RECOVERY_MAX=2
# ECAA_HARNESS_CONTRACT_ADVISORY=0
# Offline end-of-run repair pass.
# - ECAA_AUTO_REPAIR — DEFAULT ON (in code; set to 0/false/no/off to opt out).
# When enabled, the harness runs the OFFLINE repair loop ONCE at its loop-exit
# convergence point (the `after.is_complete()` block in main.rs), which
# BOTH run paths reach: the standalone/CLI run (no --session-id) and the
# session/web-UI run (the server spawns this same harness WITH a
# --session-id as the execution engine). The repair call sits OUTSIDE the
# `progress.is_none()` gate that scopes the standalone self-finalize, so it
# fires on both paths; the repair loop re-runs finalize internally and is
# idempotent. It auto-corrects deterministic prose-vs-table counts and
# re-seals BagIt manifests in place, and routes any agentic /
# offline-unverifiable gap to the signed review list
# (runtime/repair-status.json — the FullyPassing/MostlyPassing/Failing
# tri-state verdict the UI can read — plus runtime/repair-requests.jsonl).
# STRICTLY best-effort: any error or panic is logged and dropped, so it
# never changes the run outcome. It NEVER re-executes an agent at
# end-of-run; agentic re-runs remain the manual
# `ecaa-workflow repair --agent` path.
ECAA_AUTO_REPAIR=1
# Bypass the harness multi-process flock (RC-22). The harness refuses
# to start when another harness holds the per-session flock at
# ~/.ecaa-workflow/locks/<session_id>.lock (exit code 2). Setting
# =1 lets a second harness attach — concurrent harnesses race on the
# dispatch WAL, so tests only.
# ECAA_HARNESS_DEBUG_ALLOW_MULTI_PROCESS=1
# Preserve per-task scratch directories (runtime/scratch/<task_id>/)
# on Completed / Failed transitions for forensic debugging. Default
# off (the cleanup runs as soon as the task reaches a terminal state).
# ECAA_SCRATCH_KEEP=1
# Per-agent-invocation guardrails honored by all three agent wrappers.
# - ECAA_AGENT_MEMORY_CAP_GB=<n>: per-agent memory ceiling. Unset = no
# cap. When set, agent-claude.sh wraps the host-path `claude`
# subprocess tree in `systemd-run --user --scope -p MemoryMax=<n>G`
# (falls back to `prlimit --as`); docker/apptainer paths add
# `--memory=<n>g`. Only the agent cgroup is reaped on breach;
# harness/server/dev UIs stay alive.
# - ECAA_AGENT_WALLCLOCK_SECS=<n>: per-agent wallclock cap. Unset =
# no cap. The harness kills the agent subprocess group at the
# deadline and flips the task to Blocked { TimeExceeded }.
# - ECAA_AGENT_CMD=<path>: override the agent executable (default
# ./scripts/agent-claude.sh). Used by integration tests +
# alternative-agent experiments.
# - ECAA_AGENT_SCRIPT=<path>: eval-adapters alias for ECAA_AGENT_CMD.
# Honoured by crates/eval-adapters/* runners; identical semantics.
# - ECAA_DISABLE_BLAS_PRELOAD=1: opt out of the libopenblas LD_PRELOAD
# interpose for sites whose R install ships a parallel libRblas
# natively (Conda forge R, MKL-built builds).
# ECAA_AGENT_MEMORY_CAP_GB=64
# ECAA_AGENT_WALLCLOCK_SECS=14400
# ECAA_AGENT_CMD=./scripts/agent-claude.sh
# ECAA_DISABLE_BLAS_PRELOAD=1
# Max retries on transient API-transport failures (socket reset / 429 /
# 502-504) inside a single agent invocation before the harness marks the
# task Blocked. Default: 2. Must be >= 1; non-numeric or < 1 clamps to 1.
# ECAA_AGENT_TRANSIENT_MAX_ATTEMPTS=2
# Dynamic per-agent resource allocation. Default ON. When active, each
# agent's ECAA_HW_VCPUS_AVAILABLE / ECAA_HW_MEMORY_GB envelope reflects
# the host's live free budget minus an overhead reserve, split across
# concurrent picks proportional to their per-stage high-water
# requirements (compute_high_water in executor/sizing.rs). Set to 0 to
# revert to the legacy "agent sees the full host" envelope.
ECAA_HW_DYNAMIC_ALLOCATION=0
# Overhead margin held back from agents for the harness, server, UI,
# and OS. The larger of (absolute, pct of total) wins per dimension.
# Defaults: 1 vCPU, 2 GB, 10%. ECAA_HW_OVERHEAD_PCT is capped at 50.
ECAA_HW_OVERHEAD_VCPUS=1
ECAA_HW_OVERHEAD_MEMORY_GB=2
ECAA_HW_OVERHEAD_PCT=10
# Fill idle host capacity. By default dynamic allocation grants each picked
# task exactly its per-stage high-water (e.g. 8 vCPU / 64 GB), which can leave a
# large host mostly idle on a serial run (ECAA_HARNESS_CONCURRENCY=1). With this
# ON, under-subscribed picks scale UP to fill the usable budget (host free minus
# overhead), so a single heavy task (e.g. Harmony/PCA on ~1M cells) gets ~all
# free cores + RAM and finishes faster; K concurrent picks split the budget
# proportionally, so it never over-subscribes. Local executor only (AWS/SLURM
# size by instance type). Set to 0 to keep exact-high-water allocation.
ECAA_HW_FILL_HEADROOM=1
# Advisory CPU/thread floor from remediation or site policy. The agent
# bootstrap applies it to BLAS/OpenMP thread env vars and container
# `--cpus`. Unset = use scheduler's recommended thread count.
# ECAA_HW_NPROC_HINT=8
# Note: ECAA_HW_VCPUS_AVAILABLE / ECAA_HW_MEMORY_GB / ECAA_HW_GPU /
# ECAA_HW_RECOMMENDED_THREADS / ECAA_HW_TOOL_THREAD_CURVES /
# ECAA_HW_TASK_RESOURCE_CLASS / ECAA_HW_CONCURRENT_PEERS_BY_CLASS /
# ECAA_HW_INTAKE_FACTS / ECAA_HW_ENV_OVERRIDES / ECAA_HW_GPU_CAPABILITY_REF
# are RUNTIME-SET by the harness onto each agent's env at dispatch.
# Setting them in your shell will be silently overwritten — only override
# when reproducing a specific allocation in isolation.
# =============================================================================
# Pilot sizing
# =============================================================================
# Runs a small pilot against representative tasks before the main
# provisioning call, then feeds the projected high-water requirements
# into local allocation, AWS instance selection, and SLURM class
# selection. AWS defaults pilot on; local/SLURM require explicit opt-in
# and SLURM consumes an existing report when present.
# Gate for the pilot pathway. Default: 1 on AWS, 0 elsewhere.
# ECAA_PILOT_ENABLED=1
# Number of pilot tasks to sample before projecting. Default: 3.
# ECAA_PILOT_TASKS=3
# Multiplier applied to the pilot's measured high-water when projecting
# the full-run shape. Default: 1.5.
# ECAA_PILOT_MULTIPLIER=1.5
# Pilot instance type override. When set, the pilot runs on this shape
# regardless of DAG contents (useful for forcing a cheap shape like
# t3.medium). Cleared after pilot release.
# ECAA_PILOT_INSTANCE=t3.medium
# Measurement interval during the pilot (seconds). Default: 5.
# ECAA_PILOT_INTERVAL_SECS=5
# Long-form aliases accepted by the pilot config loader. Short names
# above remain canonical.
# ECAA_PILOT_INSTANCE_TYPE ≡ ECAA_PILOT_INSTANCE
# ECAA_PILOT_TASK_COUNT ≡ ECAA_PILOT_TASKS
# ECAA_PILOT_PROJECTION_MULT ≡ ECAA_PILOT_MULTIPLIER
# ECAA_PILOT_MEASUREMENT_INTERVAL_SECS ≡ ECAA_PILOT_INTERVAL_SECS
# ECAA_PILOT_INSTANCE_TYPE=t3.medium
# ECAA_PILOT_TASK_COUNT=3
# ECAA_PILOT_PROJECTION_MULT=1.5
# ECAA_PILOT_MEASUREMENT_INTERVAL_SECS=5
# =============================================================================
# Stall monitor — opt-in (default on for AWS; off elsewhere)
# =============================================================================
# Samples CPU/memory/GPU from /proc/<pid>/ (local) or CloudWatch (AWS)
# and emits a StallSignal when a threshold is breached. The harness
# translates signals into Blocked { Stalled } sessions.
# Gate for the stall monitor. Defaults to on when ECAA_EXECUTOR_MODE=aws,
# off otherwise. Accepts 1/0/true/false.
# ECAA_STALL_ENABLED=1
# CPU below this percentage counts as starvation. Default: 5.0.
# ECAA_STALL_CPU_MIN_PCT=5.0
# Sustained CPU-starvation duration before a stall signal fires (minutes).
# Default: 30.
# ECAA_STALL_CPU_WINDOW_MINS=30
# Memory above this percentage counts as pressure. Default: 90.0.
# ECAA_STALL_MEM_MAX_PCT=90.0
# Sustained memory-pressure duration before a stall signal fires (minutes).
# Default: 5.
# ECAA_STALL_MEM_WINDOW_MINS=5
# GPU-idle threshold while the envelope declares training (minutes).
# Default: 15.
# ECAA_STALL_GPU_IDLE_MINS=15
# Multiplier on expected task runtime before a runtime-overrun signal
# fires. Default: 3.0.
# ECAA_STALL_RUNTIME_MULT=3.0
# Sampling cadence (seconds). Default: 30. CloudWatch native granularity
# caps at 60s.
# ECAA_STALL_SAMPLE_INTERVAL_SECS=30
# =============================================================================
# Chat server → harness dispatch
# =============================================================================
# URL the harness POSTs progress events back to. Must match the chat
# server's actual listening port (the server's `--port` flag, default
# 3000). A mismatch silently breaks the UI's harness-progress events
# because the SSE bus never receives the POSTs (the harness logs
# `Connection refused` after retries) — WORKFLOW.json still updates,
# but the "running" pill in the inspector never lights up.
#
# The canonical dev port is 3000 across:
# - `crates/server/src/lib.rs` `--port` default
# - Makefile `DEV_SERVER_PORT`
# - `ui/vite.config.ts` VITE_API_PORT fallback
# The `dev_port_drift` test (crates/server/tests/dev_port_drift.rs)
# fails CI if these three drift apart. Default: http://127.0.0.1:3000.
ECAA_SERVER_URL=http://127.0.0.1:3000
# Vite reads this and uses it as the `/api/*` proxy target for the
# dev UI server. Must match ECAA_SERVER_URL's port above; the
# Makefile's `make up` recipe exports it automatically to keep the
# vite proxy and the server's `--port` aligned. Default: 3000.
VITE_API_PORT=3000
# Per-IP global rate-limit (tower-governor). Production defaults from
# crates/server/src/lib.rs are 2000ms interval / burst 30 → 0.5 req/s
# sustained — fine for proxied prod traffic where each SME is a
# distinct peer IP. In LOCAL DEV every UI session shares 127.0.0.1
# and the UI polls /dag, /proposals, /dispositions, /compose-outcome,
# /metrics every few seconds; that burns through the 30-token burst
# within a single chat turn and 429s the polling endpoints. The
# source comment at crates/server/src/lib.rs:240 explicitly
# recommends these dev overrides.
# ECAA_RATE_LIMIT_INTERVAL_MS=10
# ECAA_RATE_LIMIT_BURST=10000
# Absolute path to the ecaa-workflow-harness binary. The chat server
# spawns this as a subprocess for post-emit execution. Default: resolved
# via `which ecaa-workflow-harness`.
# ECAA_HARNESS_BIN_PATH=/usr/local/bin/ecaa-workflow-harness
# Default agent script the chat server passes as --agent when spawning
# the harness. Operators set this to the path of agent-claude.sh /
# agent-claude-aws.sh / agent-claude-slurm.sh. Default: ./scripts/agent-claude.sh
# ECAA_DEFAULT_AGENT_PATH=./scripts/agent-claude.sh
# Override the agent-scripts root for path-jail validation. Default
# ./scripts/. The server refuses to spawn an --agent that doesn't
# resolve under this directory.
# ECAA_SCRIPTS_DIR=./scripts
# Server-side default for harness `--max-iterations` when the chat
# server spawns it via /start-execution. Default: 20. Bump to e.g.
# 1000 for long-running detached compute so the harness lives across
# the workload instead of churning at iteration boundaries. The
# standalone `make ivd-execute` Make target reads its own MAX_ITER
# knob (default 30) and does NOT honor this variable.
ECAA_DEFAULT_MAX_ITERATIONS=20
# =============================================================================
# SLURM provisioning — required when ECAA_EXECUTOR_MODE=slurm
# =============================================================================
# The required SLURM variables are documented in this section. Every var below
# is no-op when ECAA_EXECUTOR_MODE=local|aws.
#
# Transport: system `ssh` + `rsync` via duct (ControlMaster multiplexing).
# Cost model: NoopCostModel — HPC fairshare/QoS quotas, not $/hr.
# Login/submit node hostname. REQUIRED.
# ECAA_SLURM_HOST=login.cluster.example.org
# Shared-FS path on the cluster where packages + wrappers get staged.
# REQUIRED. Typical: /scratch/$USER/scripps
# ECAA_SLURM_STAGING_DIR=/scratch/alan/scripps
# Fallback partition when the sizing mapping has nothing that fits.
# REQUIRED. Also used when the operator doesn't ship a slurm-mapping.yaml.
# ECAA_SLURM_DEFAULT_PARTITION=normal
# SSH user. Defaults to the shell's `$USER` via `ssh`'s own fallback.
# ECAA_SLURM_USER=alan
# Private key path. Omit to let `ssh` use ssh-agent / default keys.
# ECAA_SLURM_SSH_KEY=~/.ssh/id_ed25519_cluster
# Bastion host (`ssh -J <host>` flag). Omit for direct SSH.
# ECAA_SLURM_PROXY_JUMP=bastion.corp.example.org
# Billing / project account (`#SBATCH --account=`). Omit when the site
# derives it from the user.
# ECAA_SLURM_ACCOUNT=lotz-lab
# Default QoS (`#SBATCH --qos=`). Overridden per-class when a
# resource_class specifies its own qos in slurm-mapping.yaml.
# ECAA_SLURM_DEFAULT_QOS=normal
# Comma-separated `module load` list prepended to every sbatch script
# body. Leave unset when the site's default environment is sufficient.
# ECAA_SLURM_MODULES=python/3.11,singularity/3.8
# `sacct` poll rate (seconds). Default: 20. Lower wastes SSH round-trips;
# higher delays staleness detection.
# ECAA_SLURM_POLL_INTERVAL_SECS=20
# Abort + scancel a job that's still PENDING after this many seconds.
# Default: 21600 (6 hours).
# ECAA_SLURM_MAX_QUEUE_WAIT_SECS=21600
# Fallback `--time=` when the sizing mapping's resolver falls back
# (no named resource_class satisfied the requirement). Default: 04:00:00.
# ECAA_SLURM_DEFAULT_TIME_LIMIT=04:00:00
# Container runtime for the SLURM agent wrapper (agent-claude-slurm.sh).
# Accepted: singularity | apptainer | podman. Unset = agent runs on the
# host's Python/R/etc. directly. Sites that enforce rootless containers
# on compute nodes should pick one here.
# ECAA_SLURM_CONTAINER_RUNTIME=singularity
# Concurrent CPU-class task budget on the SLURM lane. Default 1 (serial).
# Bump to fan out independent CPU tasks across multiple sbatch
# submissions in parallel.
# ECAA_SLURM_MAX_PARALLEL_TASKS=1
# Concurrent GPU-class task budget. Default 0 (refuse GPU tasks on the
# SLURM lane entirely). Set to a positive value to permit GPU dispatch
# up to that ceiling; the resource_class → partition mapping must
# also resolve.
# ECAA_SLURM_MAX_GPU_TASKS=0
# Emit per-job sacct usage stats on completion (cputime, MaxRSS, etc.)
# into runtime/slurm-usage.json. Default off; opt in for capacity
# planning runs.
# ECAA_SLURM_EMIT_USAGE=1
# =============================================================================
# Container mode (Stage 15) — local agent
# =============================================================================
# Operator-level default container image used by scripts/agent-claude.sh
# when neither the per-task `container.image` field nor the package's
# policies/container.json::image is set. Per-task and package-level
# pins still win — this is the lowest-priority source. With this set,
# every agent invocation runs inside the named image by default; unset
# (or set to empty) to fall back to host execution.
#
# `bio-min:local` (built by scripts/build-bio-min.sh) ships R 4.5 +
# Python 3.11 + numpy/pandas/scipy + OpenBLAS + OpenMP + TBB + OpenMPI
# + samtools/bcftools/STAR/salmon/GATK4/fastqc + Node + claude-code CLI
# so the dispatched agent can `apt install` / `pip install` /
# `Rscript -e 'BiocManager::install("…")'` whatever analysis package
# the SME-confirmed method requires at runtime.
# The older `ecaa-workflow-agent:local` baseline shipped only
# Python+pip+node and forced agents to reimplement DESeq2 etc. in
# pure Python (2026-05-24 S1 finding).
# ECAA_DEFAULT_CONTAINER_IMAGE=bio-min:local
# Replay image-drift fallback. When an agent-free `replay` cannot find the exact
# recorded per-task snapshot image locally (e.g. it was garbage-collected), fall
# back to the current base image (ECAA_DEFAULT_CONTAINER_IMAGE, default
# `bio-min:local`) so old packages remain re-executable — logging a loud IMAGE
# DRIFT warning and tagging the report's `env_tier`. Enabled by default; set to
# 0 to REQUIRE the exact recorded image (fail loudly instead of drifting).
# ECAA_REPLAY_IMAGE_FALLBACK=1