Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
marcsun13
/
gguf-kernels
like
0
GGUF
kernel
quantization
License:
mit
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
gguf-kernels
217 MB
Ctrl+K
Ctrl+K
1 contributor
History:
12 commits
marcsun13
HF Staff
Minimal README: ops, devices, source and how to refresh it
fd5d3aa
37 minutes ago
build
Refresh the Metal builds after the command-buffer fix
about 1 hour ago
gguf_cuda
GGUF kernels: dequantize + fused gemv over packed blocks, 9 CUDA variants
2 days ago
gguf_metal
Encode on the MPS stream's own queue
about 2 hours ago
tests
GGUF kernels: dequantize + fused gemv over packed blocks, 9 CUDA variants
2 days ago
torch-ext
GGUF kernels: dequantize + fused gemv over packed blocks, 9 CUDA variants
2 days ago
vendor
Add Metal backend; vendor whole llama.cpp trees
about 22 hours ago
.gitattributes
3.63 kB
Refresh the Metal builds after the command-buffer fix
about 1 hour ago
.gitignore
95 Bytes
Minimal README: ops, devices, source and how to refresh it
37 minutes ago
README.md
1.69 kB
Minimal README: ops, devices, source and how to refresh it
37 minutes ago
SKILL.md
Safe
8.62 kB
GGUF kernels: dequantize + fused gemv over packed blocks
2 days ago
build.toml
3.43 kB
Add Metal backend; vendor whole llama.cpp trees
about 22 hours ago
flake.lock
3.05 kB
Add Metal backend; vendor whole llama.cpp trees
about 22 hours ago
flake.nix
Safe
290 Bytes
GGUF kernels: dequantize + fused gemv over packed blocks
2 days ago
vendor.py
3.29 kB
Add Metal backend; vendor whole llama.cpp trees
about 22 hours ago