14 Commits
Author SHA1 Message Date
isHuangXin 425fbbcb75 Update submodule to release-bitnet-embedding-0.6b-270m branch 2026-07-15 10:30:33 +02:00
isHuangXin 1f6a25aeeb Update llama.cpp submodule to include SafetensorRemote utility classes 2026-07-15 08:47:54 +02:00
isHuangXin 3b04140a54 feat: add I2_S GGUF conversion for bitnet-b1.58-2B-4T and refactor T-MAC LUT path
- Add quantize_to_i2_s() for direct ternary-to-I2_S packing in conversion script
- Support offline-quantized models (uint8 packed weights + weight_scale)
- Fix weight_quant double-quantization bug for offline-quantized models
- Fix I2_S scale computation to use first nonzero absolute value
- Add I2_S ftype mapping and BitNetForCausalLM registration
- Refactor ggml-bitnet-lut T-MAC wrapper with proper mul_mat implementation
- Update llama.cpp submodule with I2_S ftype and 2B model type support
2026-07-13 06:26:16 +02:00
isHuangXin 69bf64d373 Add bitnet-embeddings-0.6b model adaptation with F16 and I2_S GGUF conversion
- Add GGUF conversion tool for bitnet-embeddings-0.6b (safetensors -> F16/I2_S GGUF)
- Add Qwen3 architecture support in llama.cpp submodule with per-projection RMSNorm
- Add I2_S ternary quantization (2-bit packed -1/0/+1) for lossless precision
- Add f16 norm weight support for correct embedding inference
- Guard bitnet-lut-kernels.h include with TL1/TL2 preprocessor checks
- Update llama.cpp submodule to dev-bitnet-embedding-0.6b branch
- Document F16 (from multilingual-e5-0.6b) and I2_S (from bitnet-embeddings-0.6b) conversion process
2026-07-12 04:10:30 +02:00
XSquirrelC 1876a3e889 [merge] submodule llama.cpp 2026-01-27 03:09:32 +00:00
deva100 112f853414 [feat] I2S kernels for weight & activation parallel on Intel & ARM machine; [feat] I2S GEMV & GEMM(llama.cpp); [feat] quantize activation & dequantize embedding(llama.cpp); [fix] compile bug: cannot define __ARM_FEATURE_DOTPROD(llama.cpp) 2025-11-19 07:35:05 +00:00
younesbelkada 765741d80b update submodule 2025-05-21 11:52:30 +04:00
junhuihe 488dc1e876 Fix model architecture name 2025-04-22 17:28:59 +08:00
potassiummmm 4f2e41a514 add support for bitnet2b_2501 model 2025-03-12 18:16:45 +08:00
potassiummmm aa39c0cdcc fix version requirement of transformers pypi package and model list for codegen 2024-12-18 17:54:23 +08:00
younesbelkada c1892d6818 updated submodule 2024-11-14 14:53:43 +00:00
potassiummmm bf11a49f11 Add support for ios platform 2024-11-11 15:13:55 +08:00
Eddie-Wang1120 c82b5e6674 update 3rdparty/llama.cpp 2024-10-18 10:08:22 +08:00
potassiummmm 6cfd8831fd initial commit 2024-10-17 21:21:10 +08:00