Skip to content
@local-inference-lab

local-inference-lab

Popular repositories Loading

  1. rtx6kpro rtx6kpro Public

    RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

    Python 841 57

  2. b12x b12x Public

    Python 175 49

  3. llm-inference-bench llm-inference-bench Public

    LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

    Python 68 11

  4. blackwell-llm-docker blackwell-llm-docker Public

    Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)

    Shell 64 16

  5. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 27 20

  6. quant-toolkit quant-toolkit Public

    Python 14 7

Repositories

Showing 10 of 14 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…