Skip to content

Repository files navigation

AdCoEvo

AdCoEvo is a co-evolving multi-agent framework for advertising delivery. Planner allocates account budgets, Optimizer adjusts ad-unit delivery strategies, and Inspector diagnoses settled results to guide the next period. The agents make operational decisions across consecutive periods and update their policies from delivery feedback. LangGraph orchestrates their decisions within each period. The framework uses AuctionNet for its simulation environment and keeps its underlying Abid bidder for per-impression bidding.

中文说明:README.zh-CN.md


1. Install

git clone https://github.com/jd-opensource/AdCoEvo.git adcoevo && cd adcoevo

# Clone the full AuctionNet simulator separately and follow its Apache-2.0 licence.
# This yields ./AuctionNet/, which is where the code looks. Nothing else to configure.
git clone https://github.com/alimama-tech/AuctionNet
# put it elsewhere? export ADCOEVO_AUCTIONNET=/path/to/AuctionNet

conda create -n adcoevo python=3.10 -y
conda activate adcoevo
pip install -r requirements.txt
# torch / vllm / peft / transformers: install for your own CUDA version

2. Provide a base model and initial role adapters

The framework does not bundle model weights or initial adapters. Provide a local causal language model and three compatible LoRA adapters, one for each role. They initialize the policies that will be updated during continuous delivery.

Inspector reward normalization requires four statistics computed from a representative set of settled trajectories before the run: mu_F, sigma_F, mu_N, and sigma_N. Keep these values fixed throughout the run, save them in a JSON file, and pass that file with --inspector-stats:

{"mu_F": 0.0, "sigma_F": 1.0, "mu_N": 0.0, "sigma_N": 1.0}

These numbers illustrate the file format only; calculate the values for your own data before running the framework.

export ADCOEVO_BASE_MODEL=/path/to/your/model
# optional, only if the tokenizer lives elsewhere
export ADCOEVO_TOKENIZER=/path/to/tokenizer

3. Choose the market profile

python scripts/coevolve.py --env-profile sparse  --method adcoevo --print-config
python scripts/coevolve.py --env-profile general --method adcoevo --print-config
Profile pValue base CPAConstraint Budget (48 slots) Conversions / period
sparse 0.0005 [60, 130] 150,000 ~1,300
general 0.005 [6, 12] 150,000 ~12,500

--env-profile switches the AuctionNet simulator's market configuration. sparse retains AuctionNet's native pValue generation base, 48 unit budgets, and CPA constraints. general adjusts those three parameter groups using the public General data. See docs/environment_profiles.md for the settings and their provenance.

4. Choose the observation window, adjustment limit, and budget

Flag Meaning Default
--window W Warm-up ticks. The first W ticks run the native budget with all HOLD (pure observation, producing the causal history); agent decisions start at tick W. 12
--quota K Maximum tCPA changes per unit and period. SET actions beyond K become forced HOLD. 8
--budget-scale Multiplier for the controlled accounts' starting budgets. 1.0
python scripts/coevolve.py --env-profile sparse --method adcoevo --window 12 --quota 8 ...

5. Choose a continuous-update method

python scripts/coevolve.py --list-methods

For methods with a learned prefix value model, Planner and Inspector use one-step reward-minus-value advantages; Optimizer computes GAE separately within each account's sequence of decision ticks.

Method Collection Advantage estimate Policy surrogate
adcoevo (default) async single rollout per-role value credit token-level bilateral ratio mask + KL/epoch gate
ppo sync single rollout per-role value credit token-level clipped PPO
cispo async single rollout per-role value credit clipped detached importance weight
gspo async single rollout per-role value credit sequence-level ratio, clipped
spo rolling window persistent EMA historical baseline ratio clip, no learned critic
srppo rolling window Monte-Carlo prefix credit terminal business credit spread over tokens

Every method collects a single factual trajectory per period. The gspo entry applies GSPO's sequence-level ratio and clipping; its advantage follows the role-wise value credit above, not group sampling — the online loop never samples multiple continuations of the same state.

6. Run

python scripts/coevolve.py \
  --env-profile general --method adcoevo \
  --window 12 --quota 8 \
  --episodes 71-120 --pv 500000 --seed 30380001 \
  --run-dir runs/demo --gpu 0 --port 8001 \
  --inspector-stats /path/to/inspector_reward_stats.json \
  --planner-init   /path/to/planner-lora \
  --optimizer-init /path/to/optimizer-lora \
  --inspector-init /path/to/inspector-lora

--print-config resolves everything without launching anything. The command executes 50 consecutive periods and schedules one role for policy updating per period in Planner → Optimizer → Inspector order. Each subsequent period uses the latest published policy versions. Per-period records are saved to runs/demo/eval/arms.jsonl; the run and policy manifests remain under runs/demo/. Aggregate ADV as the mean of recorded F values and CPA as sum(C) / sum(N), rather than averaging period-level CPA values. See docs/reproduction.md for the evaluation protocol.

Within each period, the collector and fixed-policy evaluator use a compiled LangGraph workflow: warm-up observation → Planner budget decisions → one Optimizer tick at a time until the period ends → Inspector review. The Optimizer node loops through the remaining ticks. The caller carries Inspector guidance into the next period and handles policy-update scheduling separately.

--gate optionally evaluates the final frozen adapters after the continuous run. Its result is a separate check, not the continuous-update result.

7. Evaluate fixed policies without updates

Use the same three-role delivery workflow with fixed adapters for every period:

python scripts/evaluate.py \
  --env-profile general --episodes 71-120 --pv 500000 --seed 30380001 \
  --window 12 --quota 8 --temperature 0 \
  --out runs/frozen_demo --gpu 0 --port 8001 \
  --planner-init   /path/to/planner-lora \
  --optimizer-init /path/to/optimizer-lora \
  --inspector-init /path/to/inspector-lora

This starts a local vLLM server using ADCOEVO_BASE_MODEL, runs Planner, Optimizer, and Inspector on each period, and carries Inspector guidance to the next period. It starts no learners and never changes the adapters. The period records, summary, and run settings are saved under runs/frozen_demo/. --resume continues an interrupted run. Inspector reward statistics are needed for policy updates, not for this fixed-policy evaluation. To use an existing OpenAI-compatible server, pass --base-url and the three --*-model names instead of local adapter paths. See docs/reproduction.md. --print-config validates the local model and adapters and shows the resolved command without starting a server or rollout.

8. Scope and dependencies

  • ADV is a proxy computed on top of AuctionNet auction output, not real booked revenue.
  • AuctionNet provides a simulated execute → settle → update loop; these runs do not use production traffic.
  • The simulator can produce a no-operation reference rollout for each period.
  • This repository does not package the full AuctionNet simulator or its original datasets. A few compatibility helpers and native simulator parameters are adapted from AuctionNet; see NOTICE for attribution and licensing.
  • Supply your own base model, compatible initial adapters, and frozen Inspector reward statistics. None are included in this repository.

9. Further detail

Document Contents
docs/environment_profiles.md profile provenance, what is fixed vs calibrated
docs/methods.md per-method mechanics, role schedule, tuning flags
docs/reproduction.md metric definitions, run protocol, resume
NOTICE copyright and third-party attribution
LICENSE Apache License 2.0 (full text)
CONTRIBUTING.md development setup and contribution guide

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages