feat: add I2_S GGUF conversion for bitnet-b1.58-2B-4T and refactor T-MAC LUT path

- Add quantize_to_i2_s() for direct ternary-to-I2_S packing in conversion script
- Support offline-quantized models (uint8 packed weights + weight_scale)
- Fix weight_quant double-quantization bug for offline-quantized models
- Fix I2_S scale computation to use first nonzero absolute value
- Add I2_S ftype mapping and BitNetForCausalLM registration
- Refactor ggml-bitnet-lut T-MAC wrapper with proper mul_mat implementation
- Update llama.cpp submodule with I2_S ftype and 2B model type support
This commit is contained in:
isHuangXin
2026-07-13 06:26:16 +02:00
parent 69bf64d373
commit 3b04140a54
6 changed files with 1395 additions and 55 deletions
+21
View File
@@ -0,0 +1,21 @@
[Kernels_0]
m = 3200
k = 8640
bm = 160
bk = 96
bmm = 32
[Kernels_1]
m = 3200
k = 3200
bm = 320
bk = 96
bmm = 32
[Kernels_2]
m = 8640
k = 3200
bm = 320
bk = 96
bmm = 32