Go to file

16337 682063abf1 feat: 改用 4-bit NF4 纯 GPU 推理，关闭 thinking 模式

- 模型加载改为 bitsandbytes 4-bit NF4 量化，device_map={"":0} 纯 GPU
- 关闭 Qwen3.5 thinking 模式 (enable_thinking=False)
- 精度从 60% 提升到 90%，推理速度 1-2 tokens/s
- GPU 显存 7.13GB/8GB，输出质量正常
- 更新所有测试结果和综合报告

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

2026-03-16 17:38:33 +08:00

docs/plans

init: 项目初始化，添加 .gitignore 和 README

2026-03-16 11:27:17 +08:00

vsp/qwen3.5-9b

feat: 改用 4-bit NF4 纯 GPU 推理，关闭 thinking 模式

2026-03-16 17:38:33 +08:00

.gitignore

fix: 修复模型加载方式，改用 FP16+CPU offload

2026-03-16 13:05:20 +08:00

README.md

init: 项目初始化，添加 .gitignore 和 README