Repository navigation
feature: Refactor the CI workflow #531
Description
Activity
- addedfeature requestNew feature or requestNew feature or requesthelp wantedExtra attention is neededExtra attention is neededgood first issueGood for newcomersGood for newcomers
on Jun 20, 2025 Hi, I think I would interested in this issue
Hi @Shaoting-Feng,
I'd love to contribute to this issue as well. Thanks for the clear breakdown! I believe this would be a great starting point for me to get involved in the project and make some contributions.
Specifically, I'm ready to start working on:
- Eliminating duplication across pipelines, refactoring Helm-related logic across pipelines.
- Migrating scripts and assets from '.github/' to 'tests/' for better modularity.
Let me know if it’s okay for me to proceed with these parts. I’m also happy to coordinate with @Dev-Arhaan if there’s a broader plan in motion.
Looking forward to contributing!
Thanks! CC @zerofishnoodles, as he’s working on more CI/CD tests.
I would like to try it!
Thanks everyone for helping! We are so glad to see more and more people willing to contribute to the project!
Right now for the e2e test, we will still use the local backend instead of mocking.
We still need to add more e2e test for different routing logics in both static backend and k8s discovery.
If possible, we also want to add more test for the unit and infra.Also there are many functionally duplicate checks, it would be good to eliminate some of them
@zerofishnoodles - Im happy to take this on if you're still looking to tackle it. I have a bunch of self-hosted runners as well, so can fork and test on my own infra as needed.
@btdeviant Hi, That will be super great! I am looking forward to see the changes!
@zerofishnoodles - Thanks for the reply and looking forward to sharing! Spent some time yesterday trying to reverse-engineer your self-hosted setup and have some questions if you don't mind:
- What is the OS and compute limitations / capabilities of the self-hosted runner? Cores, mem, GPU?
- Is minikube a hard requirement? Are y'all open to microk8s as a potential alternative? (a bit more transparent and easier to bootstrap things like nvidia-operator)
I ask because, assuming you have sufficient compute, it opens up some cool opportunities via runner groups and parallelization in addition to having some neat, clean patterns. Im happy to put together a proposal if you'd like.
Example of the helm tests. Build once, load / push the operator to the microk8s (or minikube) cluster running on the host, then run each test in a dedicated namespace in the cluster - this would require:
- MIG capable GPU
nactions-runners running on the host
https://github.com/Mad-Deecent/production-stack/actions/runs/16407595700
Note: These are failing because my CUDA version is 12.2 in the VM I had going for the self-hosted runner, below the 12.8 requirement for the llvm-openai containers y'all had defined in the helm values - easily worked around. Also only stood up a single runner on the machine, another easily solvable thing.
@btdeviant Ofc, I am happy to share. The self-hosted runner capabilities:
OS: ubuntu 22.04
CPU: 28 (logical)
Mem: 200 GB
GPU: 2*A6000
Driver and Cuda: 550.127.05 + 12.4As to the minikube requirements, we are open to alternatives, but the script we used to setup the env is based on minikube, if using other, is it possible for you to provide a env setup script too? Also the existing ci test workflow is based on minikube too. It might be a big change.
Reacted by Shaoting and btdeviant@zerofishnoodles Thank you - very helpful! Plenty of compute to do some fun stuff and make CICD a breeze. And yeah, absolutely, will share everything for review. I'll take a peek in the other repos to get a feel for usage and patterns there too
Hey @zerofishnoodles - thanks for your patience on this. Had to put this on hold, life popped up and was trying to hack this into existence using a couple of Tesla P40's which didn't vibe well with vllm min compute requirements. Upgraded to a couple RTX Ada 5000's, still waiting on mobo and CPU to arrive and can circle back after I get that all tossed together.
Reacted by Rui ZhangHi! I'd like to pick this up. I noticed #924 attempted all four items in a single PR and was closed, and the follow-up volunteer stalled — so I'd like to try the opposite approach: small, single-purpose PRs.
I checked the current state of the repo to see what's actually still outstanding:
- Move
.github/scripts/assets →tests/— still pending:.github/currently holds 12 loose files (curl-*.sh,values-*.yaml,port-forward.sh,template-chatml.jinja), referenced from bothfunctionality-helm-chart.yml(e.g.helm install vllm ./helm -f .github/values-05-secure-vllm.yaml) androuter-e2e-test.yml(.github/template-chatml.jinja), whiletests/only containse2e/. - Fake OpenAI server for local-backend jobs — partially done: the perf-test job in
router-e2e-test.ymlalready starts mock servers, but thestatic-discovery-e2e-testjob still launches two realvllm serve facebook/opt-125mprocesses with a readiness wait (the ~2-minute bootstrap this issue wants to eliminate). - Consolidate local-backend jobs / single checkout — still pending:
router-e2e-test.ymlchecks out the repo in three separate jobs. - De-duplicate Helm launches between the k8s-discovery job in
router-e2e-test.ymlandfunctionality-helm-chart.yml— appears still pending.
Proposed plan:
- PR 1: item 1 — mechanical move of scripts/assets into
tests/+ reference updates in both workflows. Low risk, easy to review, unblocks the rest. - PR 2: item 2 for the static-discovery job — swap the real vLLM backends for the mock OpenAI server and drop the readiness wait.
- PR 3: items 3–4 — job consolidation and Helm de-duplication, scoped after seeing how PRs 1–2 land.
@Shaoting-Feng @ruizhang0101 — does this ordering work for you? Happy to start on PR 1 right away, or adjust if you'd rather see a different item first.
- Move
@ighutake-debug Yes, this ordering sounds good to me. Feel free to drop the PR.
For the e2e test, we want to have cases with true servers.Reacted by ighutake-debug- added a commit that references this issue
on Jul 30, 2026 #531 item 2 — picking up next
#1024 (item 1: move
.github/scripts/assets →tests/) merged in #1024.Opening a PR for item 2: use the existing mock OpenAI servers for the
static-discovery-e2e-testjob instead of two realvllm serveprocesses (removes ~2 min GPU bootstrap on that job). k8s-discovery and Helm functionality jobs stay on real backends, per @ruizhang0101's note that we should keep true-server e2e coverage where it matters.
Describe the feature
At present, both the functionality pipeline and the router end-to-end pipeline include several redundant jobs and steps. To streamline our CI pipeline, consider the following refactoring suggestions:
.github/totests/.Why do you need this feature?
Additional context
No response