{/* This page is auto-generated from the skill's SKILL.md by website/scripts/generate-skill-docs.py. Edit the source SKILL.md, not this page. */}

Llama Cpp

llama.cpp 本地 GGUF 推理 + HF Hub 模型发现。

Skill 元数据

来源可选 — 通过 hermes skills install official/mlops/llama-cpp 安装
路径optional-skills/mlops/inference/llama-cpp
版本2.1.2
作者Orchestra Research
许可证MIT
依赖项llama-cpp-python>=0.2.0
平台linux, macos, windows
标签llama.cpp, GGUF, Quantization, Hugging Face Hub, CPU Inference, Apple Silicon, Edge Deployment, AMD GPUs, Intel GPUs, NVIDIA, URL-first

参考:完整 SKILL.md

INFO

以下是 Hermes 在触发此 skill 时加载的完整 skill 定义。这是 skill 激活时 agent 所看到的指令内容。

llama.cpp + GGUF

当需要本地 GGUF 推理、量化选型,或为 llama.cpp 发现 Hugging Face 仓库时使用此 skill。

何时使用

  • 在 CPU、Apple Silicon、CUDA、ROCm 或 Intel GPU 上运行本地模型
  • 为某个特定 Hugging Face 仓库挑选合适的 GGUF
  • 从 Hub 构建 llama-server 或 llama-cli 命令
  • 在 Hub 上搜索已支持 llama.cpp 的模型
  • 枚举某个仓库可用的 .gguf 文件及大小
  • 针对用户的 RAM 或 VRAM 在 Q4/Q5/Q6/IQ 变体之间做选择

模型发现工作流

在请求 hf、Python 或自定义脚本之前,优先使用 URL 工作流。

  1. 在 Hub 上搜索候选仓库:
    • 基础:https://huggingface.co/models?apps=llama.cpp&sort=trending
    • 加上 search=<term> 指定模型家族
    • 当用户有尺寸限制时,加上 num_parameters=min:0,max:24B 或类似条件
  2. 用 llama.cpp 本地应用视图打开仓库:
    • https://huggingface.co/<repo>?local-app=llama.cpp
  3. 当 local-app 片段可见时,以它为准:
    • 复制确切的 llama-server 或 llama-cli 命令
    • 完全按 HF 所示报告推荐的量化
  4. 将同一个 ?local-app=llama.cpp URL 作为页面文本或 HTML 读取,提取 Hardware compatibility 下的章节:
    • 优先采用其确切的量化标签和大小,而非通用表格
    • 保留仓库特定的标签,如 UD-Q4_K_M 或 IQ4_NL_XL
    • 如果抓取到的页面源码中没有该章节,请明确说明,并回退到 tree API 加通用量化指引
  5. 查询 tree API 确认实际存在的文件:
    • https://huggingface.co/api/models/<repo>/tree/main?recursive=true
    • 保留 type 为 file 且 path 以 .gguf 结尾的条目
    • 以 path 和 size 作为文件名与字节大小的依据
    • 将量化 checkpoint 与 mmproj-*.gguf 投影器文件、BF16/ 分片文件区分开
    • 仅把 https://huggingface.co/<repo>/tree/main 作为人工回退
  6. 如果 local-app 片段在文本中不可见,根据仓库和所选量化重建命令:
    • 简写量化选择:llama-server -hf <repo>:<QUANT>
    • 精确文件回退:llama-server --hf-repo <repo> --hf-file <filename.gguf>
  7. 仅当仓库尚未提供 GGUF 文件时,才建议从 Transformers 权重转换。

快速开始

安装 llama.cpp

# macOS / Linux(最简单)
brew install llama.cpp
winget install llama.cpp
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release

直接从 Hugging Face Hub 运行

llama-cli -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q8_0
llama-server -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q8_0

运行 Hub 上的确切 GGUF 文件

当 tree API 显示自定义文件命名、或缺少确切的 HF 片段时使用。

llama-server \
    --hf-repo microsoft/Phi-3-mini-4k-instruct-gguf \
    --hf-file Phi-3-mini-4k-instruct-q4.gguf \
    -c 4096

OpenAI 兼容服务器检查

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      {"role": "user", "content": "Write a limerick about Python exceptions"}
    ]
  }'

Python 绑定(llama-cpp-python)

pip install llama-cpp-python(CUDA:CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --force-reinstall --no-cache-dir;Metal:CMAKE_ARGS="-DGGML_METAL=on" ...)。

基本生成

from llama_cpp import Llama

llm = Llama(
    model_path="./model-q4_k_m.gguf",
    n_ctx=4096,
    n_gpu_layers=35,     # 0 表示 CPU,99 表示全部卸载到 GPU
    n_threads=8,
)

out = llm("What is machine learning?", max_tokens=256, temperature=0.7)
print(out["choices"][0]["text"])

对话 + 流式

llm = Llama(
    model_path="./model-q4_k_m.gguf",
    n_ctx=4096,
    n_gpu_layers=35,
    chat_format="llama-3",   # 或 "chatml"、"mistral" 等
)

resp = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is Python?"},
    ],
    max_tokens=256,
)
print(resp["choices"][0]["message"]["content"])

# 流式
for chunk in llm("Explain quantum computing:", max_tokens=256, stream=True):
    print(chunk["choices"][0]["text"], end="", flush=True)

Embedding

llm = Llama(model_path="./model-q4_k_m.gguf", embedding=True, n_gpu_layers=35)
vec = llm.embed("This is a test sentence.")
print(f"Embedding dimension: {len(vec)}")

也可以直接从 Hub 加载 GGUF:

llm = Llama.from_pretrained(
    repo_id="bartowski/Llama-3.2-3B-Instruct-GGUF",
    filename="*Q4_K_M.gguf",
    n_gpu_layers=35,
)

选择量化

先看 Hub 页面,再用通用启发式。

  • 优先选择 HF 针对用户硬件配置标记为兼容的确切量化。
  • 通用对话从 Q4_K_M 开始。
  • 代码或技术工作,若内存允许,优先 Q5_K_M 或 Q6_K。
  • 内存预算非常紧张时,仅当用户明确把“装得下”置于质量之上,才考虑 Q3_K_M、IQ 变体或 Q2 变体。
  • 多模态仓库需单独说明 mmproj-*.gguf。投影器不是主模型文件。
  • 不要规范化仓库原生标签。页面写 UD-Q4_K_M,就报告 UD-Q4_K_M。

从仓库提取可用 GGUF

当用户问存在哪些 GGUF 时,返回:

  • 文件名
  • 文件大小
  • 量化标签
  • 是主模型还是辅助投影器

除非被要求,否则忽略:

  • README
  • BF16 分片文件
  • imatrix 数据块或校准产物

此步骤使用 tree API:

  • https://huggingface.co/api/models/<repo>/tree/main?recursive=true

对于像 unsloth/Qwen3.6-35B-A3B-GGUF 这样的仓库,local-app 页面会显示 UD-Q4_K_M、UD-Q5_K_M、UD-Q6_K、Q8_0 等量化标签,而 tree API 会暴露确切文件路径,如 Qwen3.6-35B-A3B-UD-Q4_K_M.gguf 和 Qwen3.6-35B-A3B-Q8_0.gguf,并带字节大小。用 tree API 把量化标签转换为确切文件名。

搜索模式

直接使用这些 URL 形式:

https://huggingface.co/models?apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending
https://huggingface.co/<repo>?local-app=llama.cpp
https://huggingface.co/api/models/<repo>/tree/main?recursive=true
https://huggingface.co/<repo>/tree/main

输出格式

回答发现类请求时,优先采用紧凑的结构化结果,例如:

Repo: <repo>
Recommended quant from HF: <label> (<size>)
llama-server: <command>
Other GGUFs:
- <filename> - <size>
- <filename> - <size>
Source URLs:
- <local-app URL>
- <tree API URL>

参考

  • hub-discovery.md - 纯 URL 的 Hugging Face 工作流、搜索模式、GGUF 提取与命令重建
  • advanced-usage.md — 投机解码、批量推理、语法约束生成、LoRA、多 GPU、自定义构建、基准脚本
  • quantization.md — 量化质量取舍、何时用 Q4/Q5/Q6/IQ、模型尺寸缩放、imatrix
  • server.md — 直接从 Hub 启动服务器、OpenAI API 端点、Docker 部署、NGINX 负载均衡、监控
  • optimization.md — CPU 线程、BLAS、GPU 卸载启发式、批处理调优、基准测试
  • troubleshooting.md — 安装/转换/量化/推理/服务器问题、Apple Silicon、调试

资源