llama.app : website + unified llama binary
#23875
Replies: 11 comments 16 replies
|
This is terrific! Thanks to everyone who contributes to this great tool ! |
|
Congrats on the achievement and the launch of the website! If |
|
When choosing a model, the size is missing. In my opinion, it would be better to either remove selector and leave the links to the HG, or add a size indication. |
|
This looks like a great step forward in terms of usability. I was wondering with the CLI, are there any plans to adjust arugments for better UX, or is the aim to just simply expose existing tools as subcommand with all existing arguments left exactly as-is? |
|
Missed opportunity to register llama.cpp |
|
Site looks good, two critiques:
|
|
Congrats 👏 🎉 |
|
Congrats! I saw a little inconsistency though: While Qwen's, Gemma's and Step's model tags show "XB MoE · YB active", GPT-OSS's model tag does not show that, making it sound more like the Qwen3.6-27B dense model. Converting it to show that it is MoE would be more accurate. Great work as always! |
|
Thank you for the website. One suggestion for improving the website would be to have an option for whether one prefers the Windows one-liner "irm https://llama.app/install.ps1 | iex" or the Linux one-liner "curl -LsSf https://llama.app/install.sh | sh" |
|
I'm currently trying to get CUDA prebuilt binaries for x64 Linux platforms and I notice some discrepancies which are kinda confusing: The GH actions do not ship CUDA compiled binaries, yet there are linux x64 CUDA enabled docker images. So I am confused with the intended usage or fetching of CUDA binaries for the linux platform. CUDA compilation from scratch on a colab instance takes an awful while, even with parallel jobs. Would be grateful for some guidance over this! Thanks. |
|
I installed the app with the install.sh script. I figured it was amazing to work with. It has not being updated recently, however. When I did llama update it appeared to download new stuff however it didn't. Is there a way to update manually? Should I use another install method to get newer versions? Thanks! |

Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Overview
We are launching an official website for llama.cpp: https://llama.app/
The main goal of the website is to provide a simple way for new users to install and run llama.cpp on their machines. The page has an installation command (one-liner) and links/instructions to popular GGUF models on the hub. The current version is a first iteration of many.
During install the cross-platform installer ships a single binary called
llama. The binary packs all the user-facing tooling of llama.cpp (i.e.llama-server,llama-cli, etc.) with a single CLI entry point. This is mostly following thegitexample. Currently we ship binaries for the major operating systems and we plan to iterate and improve the packaging pipeline.The webpage will also provide helpful instructions for running
llamain common use cases: chat, agentic coding, etc. These will be combined with current FOTM models from various quantization providers (e.g. Unsloth, Bartowski, etc.). There will be guidelines for integration with 3rd-party agents, creating or finding the best configuration for your device and tips for utilizing advanced llama.cpp features. We will be iterating on this and would love to hear any feedback from the community on how to improve.llamaapp is here: https://github.com/ggml-org/llama.cpp/tree/master/appHF Team: @ggml-org/hf
All reactions