Internal Comms
A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use …
Search the Hugging Face Hub for llama.cpp-compatible GGUF repos, choose the right quant, and launch the model with llama-cli or llama-server.
apps=llama.cpp.https://huggingface.co/<repo>?local-app=llama.cpp..gguf filenames with https://huggingface.co/api/models/<repo>/tree/main?recursive=true.llama-cli -hf <repo>:<QUANT> or llama-server -hf <repo>:<QUANT>.--hf-repo plus --hf-file when the repo uses custom file naming.brew install llama.cpp
winget install llama.cpp
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
make
hf auth login
https://huggingface.co/models?apps=llama.cpp&sort=trending
https://huggingface.co/models?search=Qwen3.6&apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending
llama-cli -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
llama-server \
--hf-repo unsloth/Qwen3.6-35B-A3B-GGUF \
--hf-file Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \
-c 4096
hf download <repo-without-gguf> --local-dir ./model-src
python convert_hf_to_gguf.py ./model-src \
--outfile model-f16.gguf \
--outtype f16
llama-quantize model-f16.gguf model-q4_k_m.gguf Q4_K_M
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer no-key" \
-d '{
"messages": [
{"role": "user", "content": "Write a limerick about exception handling"}
]
}'
?local-app=llama.cpp page.UD-Q4_K_M instead of normalizing them.Q4_K_M unless the repo page or hardware profile suggests otherwise.Q5_K_M or Q6_K for code or technical workloads when memory allows.Q3_K_M, Q4_K_S, or repo-specific IQ / UD-* variants for tighter RAM or VRAM budgets.mmproj-*.gguf files as projector weights, not the main checkpoint.imatrix.https://github.com/ggml-org/llama.cpphttps://huggingface.co/docs/hub/gguf-llamacpphttps://huggingface.co/docs/hub/main/local-appshttps://huggingface.co/docs/hub/agents-localhttps://huggingface.co/spaces/ggml-org/gguf-my-repoSource: Hugging Face · Apache-2.0 · SHA-256 shown alongside the download.
License file included. A license and checksum are not a security certification. Review package instructions and scripts before running them.
huggingface-local-models/SKILL.md3780 byteshuggingface-local-models/SOURCE.txt200 bytesSource and packaging checks recorded on 2026-10-03. These notes are not safety certification or measured task performance.
llama.cpp; suitable RAM/VRAM and disk space; network access for model downloads.
Model downloads can be large. Each model has separate licensing and access conditions; hardware fit is not guaranteed.
Upstream commit: ca0325bb20b2d0a1b2efa893670c4c72f79e707b
Runtime status: not tested by this catalog. Configure your client and test the skill in your own environment.
Records are supplied by the site administrator and bound to a specific package. They are not third-party safety certification. This page does not execute skills.
No published scenario records yet. Resource availability and download counts do not imply measured task performance.
Be the first to share your experience.
Sign in to leave a review →A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use …
Help address review/issue comments on the open GitHub PR for the current branch using gh CLI; verify gh auth first and prompt the user to au…
Read Hugging Face paper pages and retrieve structured links to models, datasets, code, and arXiv papers.
Compare reliable sources, track uncertainty, and turn findings into a concise decision brief.