Skip to content

feature: Refactor the CI workflow #531

Description

@Shaoting-Feng

Describe the feature

At present, both the functionality pipeline and the router end-to-end pipeline include several redundant jobs and steps. To streamline our CI pipeline, consider the following refactoring suggestions:

  • Use the fake openai server for any jobs that require a local backend to eliminate the two-minute bootstrap overhead.
  • Consolidate all local-backend jobs in the router end-to-end pipeline so the router source code only needs to be checked out once.
  • Eliminate duplication across pipelines: currently, some router end-to-end jobs launch pods using Helm templates—duplicating what the Helm functionality jobs already do. We should refactor to a more modular design with a clearer structure.
  • Move the scripts and assets from .github/ to tests/.

Why do you need this feature?

  • Provide a clearer structure so developers can understand the results better.
  • Adopt a more modular design to facilitate future expansions.
  • Minimise the time developers spend waiting for test outcomes.
  • Cut down overall CI run time, given GitHub Actions’ execution limits.

Additional context

No response

Activity

  1. Dev-Arhaan commented on Jun 24, 2025

    @Dev-Arhaan

    Hi, I think I would interested in this issue

  2. QichengZhu commented on Jun 24, 2025

    @QichengZhu

    Hi @Shaoting-Feng,

    I'd love to contribute to this issue as well. Thanks for the clear breakdown! I believe this would be a great starting point for me to get involved in the project and make some contributions.

    Specifically, I'm ready to start working on:

    • Eliminating duplication across pipelines, refactoring Helm-related logic across pipelines.
    • Migrating scripts and assets from '.github/' to 'tests/' for better modularity.

    Let me know if it’s okay for me to proceed with these parts. I’m also happy to coordinate with @Dev-Arhaan if there’s a broader plan in motion.

    Looking forward to contributing!

  3. Shaoting-Feng commented on Jun 24, 2025

    @Shaoting-Feng
    CollaboratorAuthor

    Thanks! CC @zerofishnoodles, as he’s working on more CI/CD tests.

  4. WeiqiangLv commented on Jun 25, 2025

    @WeiqiangLv

    I would like to try it!

  5. ruizhang0101 commented on Jun 26, 2025

    @ruizhang0101
    Collaborator

    Thanks everyone for helping! We are so glad to see more and more people willing to contribute to the project!

    Right now for the e2e test, we will still use the local backend instead of mocking.
    We still need to add more e2e test for different routing logics in both static backend and k8s discovery.
    If possible, we also want to add more test for the unit and infra.

    Also there are many functionally duplicate checks, it would be good to eliminate some of them

  6. btdeviant commented on Jul 20, 2025

    @btdeviant

    @zerofishnoodles - Im happy to take this on if you're still looking to tackle it. I have a bunch of self-hosted runners as well, so can fork and test on my own infra as needed.

  7. ruizhang0101 commented on Jul 21, 2025

    @ruizhang0101
    Collaborator

    @btdeviant Hi, That will be super great! I am looking forward to see the changes!

  8. btdeviant commented on Jul 21, 2025

    @btdeviant

    @zerofishnoodles - Thanks for the reply and looking forward to sharing! Spent some time yesterday trying to reverse-engineer your self-hosted setup and have some questions if you don't mind:

    1. What is the OS and compute limitations / capabilities of the self-hosted runner? Cores, mem, GPU?
    2. Is minikube a hard requirement? Are y'all open to microk8s as a potential alternative? (a bit more transparent and easier to bootstrap things like nvidia-operator)

    I ask because, assuming you have sufficient compute, it opens up some cool opportunities via runner groups and parallelization in addition to having some neat, clean patterns. Im happy to put together a proposal if you'd like.

    Example of the helm tests. Build once, load / push the operator to the microk8s (or minikube) cluster running on the host, then run each test in a dedicated namespace in the cluster - this would require:

    • MIG capable GPU
    • n actions-runners running on the host

    https://github.com/Mad-Deecent/production-stack/actions/runs/16407595700

    Note: These are failing because my CUDA version is 12.2 in the VM I had going for the self-hosted runner, below the 12.8 requirement for the llvm-openai containers y'all had defined in the helm values - easily worked around. Also only stood up a single runner on the machine, another easily solvable thing.

  9. ruizhang0101 commented on Jul 21, 2025

    @ruizhang0101
    Collaborator

    @btdeviant Ofc, I am happy to share. The self-hosted runner capabilities:

    OS: ubuntu 22.04
    CPU: 28 (logical)
    Mem: 200 GB
    GPU: 2*A6000
    Driver and Cuda: 550.127.05 + 12.4

    As to the minikube requirements, we are open to alternatives, but the script we used to setup the env is based on minikube, if using other, is it possible for you to provide a env setup script too? Also the existing ci test workflow is based on minikube too. It might be a big change.

  10. btdeviant commented on Jul 21, 2025

    @btdeviant

    @zerofishnoodles Thank you - very helpful! Plenty of compute to do some fun stuff and make CICD a breeze. And yeah, absolutely, will share everything for review. I'll take a peek in the other repos to get a feel for usage and patterns there too

  11. btdeviant commented on Nov 9, 2025

    @btdeviant

    Hey @zerofishnoodles - thanks for your patience on this. Had to put this on hold, life popped up and was trying to hack this into existence using a couple of Tesla P40's which didn't vibe well with vllm min compute requirements. Upgraded to a couple RTX Ada 5000's, still waiting on mobo and CPU to arrive and can circle back after I get that all tossed together.

  12. ighutake-debug commented on Jul 27, 2026

    @ighutake-debug
    Contributor

    Hi! I'd like to pick this up. I noticed #924 attempted all four items in a single PR and was closed, and the follow-up volunteer stalled — so I'd like to try the opposite approach: small, single-purpose PRs.

    I checked the current state of the repo to see what's actually still outstanding:

    1. Move .github/ scripts/assets → tests/ — still pending: .github/ currently holds 12 loose files (curl-*.sh, values-*.yaml, port-forward.sh, template-chatml.jinja), referenced from both functionality-helm-chart.yml (e.g. helm install vllm ./helm -f .github/values-05-secure-vllm.yaml) and router-e2e-test.yml (.github/template-chatml.jinja), while tests/ only contains e2e/.
    2. Fake OpenAI server for local-backend jobs — partially done: the perf-test job in router-e2e-test.yml already starts mock servers, but the static-discovery-e2e-test job still launches two real vllm serve facebook/opt-125m processes with a readiness wait (the ~2-minute bootstrap this issue wants to eliminate).
    3. Consolidate local-backend jobs / single checkout — still pending: router-e2e-test.yml checks out the repo in three separate jobs.
    4. De-duplicate Helm launches between the k8s-discovery job in router-e2e-test.yml and functionality-helm-chart.yml — appears still pending.

    Proposed plan:

    • PR 1: item 1 — mechanical move of scripts/assets into tests/ + reference updates in both workflows. Low risk, easy to review, unblocks the rest.
    • PR 2: item 2 for the static-discovery job — swap the real vLLM backends for the mock OpenAI server and drop the readiness wait.
    • PR 3: items 3–4 — job consolidation and Helm de-duplication, scoped after seeing how PRs 1–2 land.

    @Shaoting-Feng @ruizhang0101 — does this ordering work for you? Happy to start on PR 1 right away, or adjust if you'd rather see a different item first.

  13. ruizhang0101 commented on Jul 28, 2026

    @ruizhang0101
    Collaborator

    @ighutake-debug Yes, this ordering sounds good to me. Feel free to drop the PR.
    For the e2e test, we want to have cases with true servers.

  14. ighutake-debug commented on Sep 13, 2026

    @ighutake-debug
    Contributor

    #531 item 2 — picking up next

    #1024 (item 1: move .github/ scripts/assets → tests/) merged in #1024.

    Opening a PR for item 2: use the existing mock OpenAI servers for the static-discovery-e2e-test job instead of two real vllm serve processes (removes ~2 min GPU bootstrap on that job). k8s-discovery and Helm functionality jobs stay on real backends, per @ruizhang0101's note that we should keep true-server e2e coverage where it matters.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions