Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
-
Updated
Aug 31, 2026 - Python
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Faster Whisper transcription with CTranslate2
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
[🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment Research Platform. 📃 online docs: https://ufund-me.github.io/Qbot ✨ :news: qbot-mini: https://github.com/Charmve/iQuant
A vector index built on TurboQuant, written in Rust with Python bindings
Accessible large language models via k-bit quantization for PyTorch.
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Lossy PNG compressor — pngquant command based on libimagequant library
An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.
Fast inference engine for Transformer models
[ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Pretrained language model and its related optimization techniques developed by Huawei Noah's Ark Lab.
Sparsity-aware deep learning inference runtime for CPUs
Your Cheat Sheet for AI Engineering Interview – Questions and Answers.
A model library for exploring state-of-the-art deep learning topologies and techniques for optimizing Natural Language Processing neural networks
To associate your repository with the quantization topic, visit your repo's landing page and select "manage topics."