Commit Graph
13 Commits
Author SHA1 Message Date
isHuangXin 3b04140a54 feat: add I2_S GGUF conversion for bitnet-b1.58-2B-4T and refactor T-MAC LUT path
- Add quantize_to_i2_s() for direct ternary-to-I2_S packing in conversion script
- Support offline-quantized models (uint8 packed weights + weight_scale)
- Fix weight_quant double-quantization bug for offline-quantized models
- Fix I2_S scale computation to use first nonzero absolute value
- Add I2_S ftype mapping and BitNetForCausalLM registration
- Refactor ggml-bitnet-lut T-MAC wrapper with proper mul_mat implementation
- Update llama.cpp submodule with I2_S ftype and 2B model type support
2026-07-13 06:26:16 +02:00
isHuangXin 69bf64d373 Add bitnet-embeddings-0.6b model adaptation with F16 and I2_S GGUF conversion
- Add GGUF conversion tool for bitnet-embeddings-0.6b (safetensors -> F16/I2_S GGUF)
- Add Qwen3 architecture support in llama.cpp submodule with per-projection RMSNorm
- Add I2_S ternary quantization (2-bit packed -1/0/+1) for lossless precision
- Add f16 norm weight support for correct embedding inference
- Guard bitnet-lut-kernels.h include with TL1/TL2 preprocessor checks
- Update llama.cpp submodule to dev-bitnet-embedding-0.6b branch
- Document F16 (from multilingual-e5-0.6b) and I2_S (from bitnet-embeddings-0.6b) conversion process
2026-07-12 04:10:30 +02:00
XSquirrelC e8c8107dcf [modify] some test picture and add power test script 2026-01-25 06:51:33 +00:00
deva100 7ea1f2601f [modify] fine_tuning_result.png 2026-01-20 07:40:37 +00:00
deva100 b68802ff17 [fix] embed-quant q6_k; [modify] README update 2026-01-20 04:56:50 +00:00
deva100 35b1c28585 [fix] correct README 2026-01-15 03:44:50 +00:00
deva100 53ffe5e92b [chore] update README 2026-01-15 03:37:16 +00:00
deva100 112f853414 [feat] I2S kernels for weight & activation parallel on Intel & ARM machine; [feat] I2S GEMV & GEMM(llama.cpp); [feat] quantize activation & dequantize embedding(llama.cpp); [fix] compile bug: cannot define __ARM_FEATURE_DOTPROD(llama.cpp) 2025-11-19 07:35:05 +00:00
potassiummmm 37c247c4dc Merge pull request #79 from MrEcco/main
Getting Errors When following Readme Instruction on ARM server with ubuntu24.04 #74
2024-11-07 00:30:10 +08:00
Goran Jelic-Cizmek 9d37b8692d Add GCC to compiler check 2024-10-25 16:22:55 +02:00
Andre Buryndin 70804c68e4 Fixing compilation error for ARM64+TL1 settings: microsoft#74 2024-10-23 21:50:04 +02:00
Yury 60766967a4 Fix memory leak in quantize_i2_s 2024-10-21 12:27:46 +03:00
potassiummmm 6cfd8831fd initial commit 2024-10-17 21:21:10 +08:00