AdCoEvo is a co-evolving multi-agent framework for advertising delivery. Planner allocates account budgets, Optimizer adjusts ad-unit delivery strategies, and Inspector diagnoses settled results to guide the next period. The agents make operational decisions across consecutive periods and update their policies from delivery feedback. LangGraph orchestrates their decisions within each period. The framework uses AuctionNet for its simulation environment and keeps its underlying Abid bidder for per-impression bidding.
中文说明:README.zh-CN.md
git clone https://github.com/jd-opensource/AdCoEvo.git adcoevo && cd adcoevo
# Clone the full AuctionNet simulator separately and follow its Apache-2.0 licence.
# This yields ./AuctionNet/, which is where the code looks. Nothing else to configure.
git clone https://github.com/alimama-tech/AuctionNet
# put it elsewhere? export ADCOEVO_AUCTIONNET=/path/to/AuctionNet
conda create -n adcoevo python=3.10 -y
conda activate adcoevo
pip install -r requirements.txt
# torch / vllm / peft / transformers: install for your own CUDA versionThe framework does not bundle model weights or initial adapters. Provide a local causal language model and three compatible LoRA adapters, one for each role. They initialize the policies that will be updated during continuous delivery.
Inspector reward normalization requires four statistics computed from a
representative set of settled trajectories before the run: mu_F, sigma_F,
mu_N, and sigma_N. Keep these values fixed throughout the run, save them
in a JSON file, and pass that file with --inspector-stats:
{"mu_F": 0.0, "sigma_F": 1.0, "mu_N": 0.0, "sigma_N": 1.0}These numbers illustrate the file format only; calculate the values for your own data before running the framework.
export ADCOEVO_BASE_MODEL=/path/to/your/model
# optional, only if the tokenizer lives elsewhere
export ADCOEVO_TOKENIZER=/path/to/tokenizerpython scripts/coevolve.py --env-profile sparse --method adcoevo --print-config
python scripts/coevolve.py --env-profile general --method adcoevo --print-config| Profile | pValue base | CPAConstraint | Budget (48 slots) | Conversions / period |
|---|---|---|---|---|
sparse |
0.0005 | [60, 130] | 150,000 | ~1,300 |
general |
0.005 | [6, 12] | 150,000 | ~12,500 |
--env-profile switches the AuctionNet simulator's market configuration.
sparse retains AuctionNet's native pValue generation base, 48 unit budgets,
and CPA constraints. general adjusts those three parameter groups using the
public General data. See docs/environment_profiles.md
for the settings and their provenance.
| Flag | Meaning | Default |
|---|---|---|
--window W |
Warm-up ticks. The first W ticks run the native budget with all HOLD (pure observation, producing the causal history); agent decisions start at tick W. |
12 |
--quota K |
Maximum tCPA changes per unit and period. SET actions beyond K become forced HOLD. |
8 |
--budget-scale |
Multiplier for the controlled accounts' starting budgets. | 1.0 |
python scripts/coevolve.py --env-profile sparse --method adcoevo --window 12 --quota 8 ...python scripts/coevolve.py --list-methodsFor methods with a learned prefix value model, Planner and Inspector use one-step reward-minus-value advantages; Optimizer computes GAE separately within each account's sequence of decision ticks.
| Method | Collection | Advantage estimate | Policy surrogate |
|---|---|---|---|
adcoevo (default) |
async single rollout | per-role value credit | token-level bilateral ratio mask + KL/epoch gate |
ppo |
sync single rollout | per-role value credit | token-level clipped PPO |
cispo |
async single rollout | per-role value credit | clipped detached importance weight |
gspo |
async single rollout | per-role value credit | sequence-level ratio, clipped |
spo |
rolling window | persistent EMA historical baseline | ratio clip, no learned critic |
srppo |
rolling window | Monte-Carlo prefix credit | terminal business credit spread over tokens |
Every method collects a single factual trajectory per period. The gspo entry applies
GSPO's sequence-level ratio and clipping; its advantage follows the role-wise
value credit above, not group sampling — the online loop never samples
multiple continuations of the same state.
python scripts/coevolve.py \
--env-profile general --method adcoevo \
--window 12 --quota 8 \
--episodes 71-120 --pv 500000 --seed 30380001 \
--run-dir runs/demo --gpu 0 --port 8001 \
--inspector-stats /path/to/inspector_reward_stats.json \
--planner-init /path/to/planner-lora \
--optimizer-init /path/to/optimizer-lora \
--inspector-init /path/to/inspector-lora--print-config resolves everything without launching anything. The command
executes 50 consecutive periods and schedules one role for policy updating
per period in Planner → Optimizer → Inspector order. Each subsequent period
uses the latest published policy versions. Per-period records are saved to
runs/demo/eval/arms.jsonl; the run and policy manifests remain under
runs/demo/. Aggregate ADV as the mean of recorded F values and CPA as
sum(C) / sum(N), rather than averaging period-level CPA values. See
docs/reproduction.md for the evaluation protocol.
Within each period, the collector and fixed-policy evaluator use a compiled LangGraph workflow: warm-up observation → Planner budget decisions → one Optimizer tick at a time until the period ends → Inspector review. The Optimizer node loops through the remaining ticks. The caller carries Inspector guidance into the next period and handles policy-update scheduling separately.
--gate optionally evaluates the final frozen adapters after the continuous
run. Its result is a separate check, not the continuous-update result.
Use the same three-role delivery workflow with fixed adapters for every period:
python scripts/evaluate.py \
--env-profile general --episodes 71-120 --pv 500000 --seed 30380001 \
--window 12 --quota 8 --temperature 0 \
--out runs/frozen_demo --gpu 0 --port 8001 \
--planner-init /path/to/planner-lora \
--optimizer-init /path/to/optimizer-lora \
--inspector-init /path/to/inspector-loraThis starts a local vLLM server using ADCOEVO_BASE_MODEL, runs Planner,
Optimizer, and Inspector on each period, and carries Inspector guidance to the
next period. It starts no learners and never changes the adapters. The period
records, summary, and run settings are saved under runs/frozen_demo/.
--resume continues an interrupted run. Inspector reward statistics are needed
for policy updates, not for this fixed-policy evaluation. To use an existing
OpenAI-compatible server, pass --base-url and the three --*-model names
instead of local adapter paths. See docs/reproduction.md.
--print-config validates the local model and adapters and shows the resolved
command without starting a server or rollout.
- ADV is a proxy computed on top of AuctionNet auction output, not real booked revenue.
- AuctionNet provides a simulated execute → settle → update loop; these runs do not use production traffic.
- The simulator can produce a no-operation reference rollout for each period.
- This repository does not package the full AuctionNet simulator or its original
datasets. A few compatibility helpers and native simulator parameters are
adapted from AuctionNet; see
NOTICEfor attribution and licensing. - Supply your own base model, compatible initial adapters, and frozen Inspector reward statistics. None are included in this repository.
| Document | Contents |
|---|---|
docs/environment_profiles.md |
profile provenance, what is fixed vs calibrated |
docs/methods.md |
per-method mechanics, role schedule, tuning flags |
docs/reproduction.md |
metric definitions, run protocol, resume |
NOTICE |
copyright and third-party attribution |
LICENSE |
Apache License 2.0 (full text) |
CONTRIBUTING.md |
development setup and contribution guide |