hiyouga/LLaMA-Factory
Unify Efficient Fine-tuning of 100+ LLMs
Language:Python
Total stars: 12204
Stars trend:
#python
#agent, #baichuan, #chatglm, #finetuning, #generativeai, #gpt, #instructiontuning, #languagemodel, #largelanguagemodels, #llama, #llm, #lora, #mistral, #mixtureofexperts, #peft, #qlora, #quantization, #qwen, #rlhf, #transformers
Unify Efficient Fine-tuning of 100+ LLMs
Language:Python
Total stars: 12204
Stars trend:
28 Feb 2024
2am ▋ +5
3am ▋ +5
4am +0
5am +0
6am ▉ +7
7am █ +8
8am █▎ +10
9am ▉ +7
10am ▎ +2
11am ▍ +3
12pm ▋ +5
1pm ▍ +3#python
#agent, #baichuan, #chatglm, #finetuning, #generativeai, #gpt, #instructiontuning, #languagemodel, #largelanguagemodels, #llama, #llm, #lora, #mistral, #mixtureofexperts, #peft, #qlora, #quantization, #qwen, #rlhf, #transformers
👍3
RahulSChand/gpu_poor
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
Language:JavaScript
Total stars: 878
Stars trend:
#javascript
#ggml, #gpu, #huggingface, #languagemodel, #llama, #llama2, #llamacpp, #llm, #pytorch, #quantization
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
Language:JavaScript
Total stars: 878
Stars trend:
5 Oct 2024
9am ▋ +5
10am ▏ +1
11am ▉ +7
12pm ▌ +4
1pm █▎ +10
2pm █▏ +9
3pm █▏ +9
4pm █▍ +11
5pm ▊ +6
6pm █▎ +10
7pm █▍ +11#javascript
#ggml, #gpu, #huggingface, #languagemodel, #llama, #llama2, #llamacpp, #llm, #pytorch, #quantization
adithya-s-k/AI-Engineering.academy
Mastering Applied AI, One Concept at a Time
Language:Jupyter Notebook
Total stars: 262
Stars trend:
#jupyternotebook
#finetuning, #finetuning, #finetuningllms, #inference, #largelanguagemodels, #llm, #python, #quantization
Mastering Applied AI, One Concept at a Time
Language:Jupyter Notebook
Total stars: 262
Stars trend:
3 Dec 2024
7pm ▎ +2
8pm ▍ +3
9pm +0
10pm ▏ +1
11pm ▍ +3
4 Dec 2024
12am ▊ +6
1am █▏ +9
2am █▉ +15
3am █▏ +9
4am █▍ +11
5am █▍ +11
6am █▍ +11#jupyternotebook
#finetuning, #finetuning, #finetuningllms, #inference, #largelanguagemodels, #llm, #python, #quantization
hiyouga/LLaMA-Factory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Language:Python
Total stars: 38120
Stars trend:
#python
#agent, #ai, #chatglm, #finetuning, #gpt, #instructiontuning, #languagemodel, #largelanguagemodels, #llama, #llama3, #llm, #lora, #mistral, #moe, #peft, #qlora, #quantization, #qwen, #rlhf, #transformers
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Language:Python
Total stars: 38120
Stars trend:
14 Jan 2025
12am ▏ +1
1am ▏ +1
2am ▊ +6
3am █▍ +11
4am ▋ +5
5am ▋ +5
6am █▏ +9
7am ██▍ +19
8am █ +8
9am █▎ +10
10am ▋ +5
11am ▊ +6#python
#agent, #ai, #chatglm, #finetuning, #gpt, #instructiontuning, #languagemodel, #largelanguagemodels, #llama, #llama3, #llm, #lora, #mistral, #moe, #peft, #qlora, #quantization, #qwen, #rlhf, #transformers
dipampaul17/KVSplit
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.
Language:Python
Total stars: 144
Stars trend:
#python
#applesilicon, #generativeai, #kvcache, #llamacpp, #llm, #m1, #m2, #m3, #memoryoptimization, #metal, #optimization, #quantization
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.
Language:Python
Total stars: 144
Stars trend:
16 May 2025
7pm ▏ +1
8pm █████▌ +44
9pm ████▊ +38
10pm ███▋ +29
11pm ██▎ +18#python
#applesilicon, #generativeai, #kvcache, #llamacpp, #llm, #m1, #m2, #m3, #memoryoptimization, #metal, #optimization, #quantization
hiyouga/LLaMA-Factory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Language:Python
Total stars: 49776
Stars trend:
#python
#agent, #ai, #chatglm, #finetuning, #gpt, #instructiontuning, #languagemodel, #largelanguagemodels, #llama, #llama3, #llm, #lora, #mistral, #moe, #peft, #qlora, #quantization, #qwen, #rlhf, #transformers
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Language:Python
Total stars: 49776
Stars trend:
24 May 2025
6am ▏ +1
7am █ +8
8am ▋ +5
9am █▎ +10
10am █▋ +13
11am ███ +24
12pm █▊ +14
1pm █▌ +12
2pm █▊ +14
3pm ██▌ +20
4pm ███▏ +25
5pm ██▏ +17#python
#agent, #ai, #chatglm, #finetuning, #gpt, #instructiontuning, #languagemodel, #largelanguagemodels, #llama, #llama3, #llm, #lora, #mistral, #moe, #peft, #qlora, #quantization, #qwen, #rlhf, #transformers
adithya-s-k/AI-Engineering.academy
Mastering Applied AI, One Concept at a Time
Language:Jupyter Notebook
Total stars: 1210
Stars trend:
#jupyternotebook
#finetuning, #finetuning, #finetuningllms, #inference, #largelanguagemodels, #llm, #python, #quantization
Mastering Applied AI, One Concept at a Time
Language:Jupyter Notebook
Total stars: 1210
Stars trend:
3 Nov 2025
8am ▏ +1
9am ▎ +2
10am ▎ +2
11am ▏ +1
12pm ▎ +2
1pm ▎ +2
2pm ▎ +2
3pm ▉ +7
4pm ▏ +1
5pm +0
6pm ▏ +1
7pm ▎ +2#jupyternotebook
#finetuning, #finetuning, #finetuningllms, #inference, #largelanguagemodels, #llm, #python, #quantization
HarryR/z80ai
Z80-μLM is a 2-bit quantized language model small enough to run on an 8-bit Z80 processor. Train conversational models in Python, export them as CP/M .COM binaries, and chat with your vintage computer.
Language:Python
Total stars: 489
Stars trend:
#python
#chatbot, #codegolf, #cpm, #languagemodel, #machinelearning, #nlp, #quantization, #retro, #retrocomputing, #tinyml, #z80
Z80-μLM is a 2-bit quantized language model small enough to run on an 8-bit Z80 processor. Train conversational models in Python, export them as CP/M .COM binaries, and chat with your vintage computer.
Language:Python
Total stars: 489
Stars trend:
29 Dec 2025
10am ███▏ +25
11am ██▉ +23
12pm █▌ +12
1pm █▌ +12
2pm █▏ +9
3pm ██▏ +17
4pm █▌ +12
5pm █▍ +11
6pm █▊ +14
7pm ▉ +7
8pm █ +8
9pm ▊ +6#python
#chatbot, #codegolf, #cpm, #languagemodel, #machinelearning, #nlp, #quantization, #retro, #retrocomputing, #tinyml, #z80
RyanCodrai/turbovec
A vector index built on TurboQuant, written in Rust with Python bindings
Language:Python
Total stars: 2066
Stars trend:
#python
#ann, #avx512, #embedding, #embeddings, #faiss, #nearestneighbor, #neon, #python, #quant, #quantization, #rag, #rust, #simd, #turboquant, #vectorsearch
A vector index built on TurboQuant, written in Rust with Python bindings
Language:Python
Total stars: 2066
Stars trend:
21 May 2026
4pm ▏ +1
5pm +0
6pm +0
7pm ▏ +1
8pm ▍ +3
9pm ▎ +2
10pm ▍ +3
11pm ▉ +7
22 May 2026
12am ▊ +6
1am █ +8#python
#ann, #avx512, #embedding, #embeddings, #faiss, #nearestneighbor, #neon, #python, #quant, #quantization, #rag, #rust, #simd, #turboquant, #vectorsearch
FareedKhan-dev/kimi-k3-in-c
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Language:C
Total stars: 3210
#c
#avx2, #c99, #cpuinference, #deeplearning, #fromscratch, #inferenceengine, #kimik3, #linearattention, #llm, #llminference, #machinelearning, #memoryefficient, #mixtureofexperts, #moe, #mxfp4, #quantization, #simd, #systemsprogramming, #transformer, #zerodependencies
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Language:C
Total stars: 3210
#c
#avx2, #c99, #cpuinference, #deeplearning, #fromscratch, #inferenceengine, #kimik3, #linearattention, #llm, #llminference, #machinelearning, #memoryefficient, #mixtureofexperts, #moe, #mxfp4, #quantization, #simd, #systemsprogramming, #transformer, #zerodependencies