PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B.The language model occupies 5.93 GB, against 53.80 GB in FP16.PrismML reports that it keeps 98.2% of the parent model’s average across 20 benchmarks.The model accepts text and images and supports a 262K-token context.
PrismML demos it driving Cline coding agents and computer use on an RTX 5090.It arrives 2 months after the first Bonsai 27B, whose ternary variant retained about 95%.Is it deployable?Yes.The Apache 2.0 weights run today on a 16 GB laptop or a single 24 GB GPU.You need PrismML’s llama.
cpp fork or its MLX runtime.What is Ternary Bonsai 2 27B?The model keeps the Qwen3.8 27B architecture unchanged.It has 27.36B parameters.That splits into a 24.35B language backbone, 2.54B in embeddings and LM head, and a 0.47B vision tower.
The backbone uses hybrid attention, with about 75% linear-attention and 25% full-attention layers.Ternary weights cover embeddings, attention projections, MLP projections and the LM head.Only 26.2M parameters, or 0.0976%, stay in higher precision.
Those are the recurrent state path and normalization weights.In GGUF, the vision tower ships separately as a 0.63 GB file, loaded only for image input.How Does the Ternary Format Work?Each weight takes 1 of 3 values: -1, 0 or +1.Every group of 128 weights shares 1 FP16 scale.
A ternary value carries log2(3), or about 1.585 bits.Adding 16 scale bits per 128 weights gives 1.71 bits per weight.Counting the high-precision tensors brings the model to 1.72.Real kernels need a packed layout, so the whitepaper describes 2 GGUF packings.PTQ10 packs trits densely at 1.
76 bits per weight and 5.93 GB.PQ20 stores each trit in a 2-bit slot at 7.25 GB, which is cheaper to unpack.Weights are also stored in a rotated basis.PrismML applies a blockwise Hadamard rotation with block size 1,024 before ternary assignment.
The runtime applies the matching transform to activations before each multiply.The whitepaper cites SpinQuant for this idea.PrismML does not publish how it assigns the ternary values.(function(){var f=document.getElementById("mtp-bonsai2-27b-x9k4");window.
addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.mtpH){f.style.height=e.data.mtpH+"px";}});})(); How Does It Score Against Qwen3.8 27B?PrismML evaluated all models in thinking mode with EvalScope and vLLM on H100 GPUs.CapabilityQwen3.6 27BQwen3.
8 27BTernary Bonsai 2 27BRetentionKnowledge and reasoning84.7186.6683.9596.9%Math94.6497.0696.5799.5%Coding82.5782.1781.5899.3%Agentic and tool calling80.0579.7477.5797.3%Instruction following74.5381.2582.66101.7%Vision79.8281.6478.5996.3%Overall (20)83.685.483.998.
2% The comparison with conventional quantization is the sharper result.An IQ2XXS build of Qwen3.8 27B averages 75.2 at 7.3 GB.On AIME26 it scores 78.6, while Bonsai 2 scores 95.83.On LiveCodeBench v6 the gap is 70.05 versus 90.07.Where Does It Still Lose Quality?The 98.
2% figure is an average, and the losses are uneven.Vision retains 96.3% and knowledge and reasoning retains 96.9%.Long-horizon agent work drops further.Bonsai 2 scores 52.8 on Terminal-Bench 2.1, against 69.7 for Qwen3.8 27B.On SWE-bench Verified it scores 60.8 against 80.6.
That is about 75% retention, and both sit outside the 20-benchmark average.Reasoning effort matters too.At medium effort the model averages 79.3, against 82.6 for the FP16 baseline.Low effort is not supported.All results are PrismML’s own and have not been independently reproduced.
How Fast is It on Real Hardware?Figures are batch size 1 decode on PrismML’s custom kernels, measured September 16, 2026.An RTX 5090 reaches 142.5 tokens per second at 0.582 mWh per token.An RTX 4090 reaches 96.7 with PTQ10, and a 72 W L4 reaches 32.1.On Apple laptops, an M5 Max reaches 46.
8 and an M5 Pro reaches 27.7.Neither packing wins everywhere.PTQ10 is faster on Ada-generation cards and the L4.PQ20 is faster on Blackwell, Hopper, Ampere and Apple silicon, and at prompt processing everywhere.
PrismML research team also claims 40% better energy efficiency than a full-precision 8B model.How Do You Run It?The GGUF files need PrismML’s llama.cpp fork.Stock llama.cpp rejects the PTQ10 and PQ20 types.The Bonsai-demo repo is the supported path.Run ./setup.sh, then ./scripts/startllamaserver.
sh for chat, vision and tools at localhost:8080.Mac users can take the MLX pack, which needs its bundled loader.A WebGPU demo runs the model inside a browser.Key Takeaways 5.93 GB language model, about 9.1x smaller than the 53.80 GB FP16 baseline.83.9 average on 20 benchmarks, versus 85.4 for Qwen3.
8 27B in FP16.142.5 tokens per second on an RTX 5090 and 46.8 on an M5 Max.Long-horizon agent benchmarks keep only about 75% of full-precision scores.Stock llama.cpp cannot load these files.PrismML’s fork is required.
Check out the Whitepaper, Model weights, GitHub repo, Docs, WebGPU demo and announcement on X.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?
now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.
8 27B Performance appeared first on MarkTechPost.
Related
相關文章

美國 AI 巨頭提議開發減速,歐洲同行和政界並不認同
作者:清源 責編:清源 評論: 9 月 18 日消息,據路透社今天(18 日)報道,此前,阿莫迪、奧爾特曼和馬斯克先後發出警告,能力不斷增強的 AI 系統可能帶來風險,故有必要控制其發展節奏,但歐洲企業和官員對此表現出明顯懷疑。法國初創企業 Mistral 在聲明中指出,這些風險早在幾個月前就已十分明確了。

MiniMax Code CLI 正式開源,在評測中取得 76.7% 的任務通過率
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 Domado 的線索投遞!9 月 18 日消息,MiniMax 今晚宣佈,MiniMax Code CLI 的 v0.4.12 版本面向全球開發者正式開放,並且以 MIT 協議正式開放源代碼。

北京發佈“詞元經濟十條”:高標準建設詞元工廠,推動關鍵核心技術攻關
作者:浩渺 責編:浩渺 評論: 感謝網友 蛋殼兒 的線索投遞!9 月 18 日消息,北京市經濟和信息化局今日宣佈,北京市“詞元經濟十條”正式發佈。為貫徹落實《國務院關於深入實施“人工智能+”行動的意見》(國發〔2025〕11 號),率先培育智能經濟新形態,以詞元(Token)為抓手,發展詞元經濟新增量,制定《北京市加快詞元經濟發展的行動方案(2026—2028 年)》(注:以下簡稱《行動方案》)。

OpenAI剛曝光循環架構,這家公司更早將其用於世界模型
鍵詞只有一個:循環。 Astra採用了一種被稱為“循環深度”(Recurrent Depth)的架構,本質是讓同一組Transformer層被反覆複用,用更少的參數實現更深的計算。 這被外界視為OpenAI對傳統“堆參數、堆算力”路線的一次重大修正——不再只是把模型做大,而是讓模型學會“反覆思考”。 消息一齣,整個AI社區迅速升溫。

千問辦公接入高德門店經營專家套件 提升實體店選址與經營效率
此項新功能自9月18日起正式上線,用戶只需在千問辦公添加相應套件並完成高德賬號授權,即可實現從門店選址到日常經營分析的全鏈路工作。隨著實體商業的不斷發展,傳統的人工選址和經營分析方式顯得效率低下且容易出錯。個體創業者在開新店時,往往需要花費大量時間進行實地勘察,整體耗時可達一週,而商家在監測門店熱度變化及競爭對手動態時,也面臨數據整理繁瑣、分析結果不準確的問題。

阿里:千問辦公協助國家天文臺科研團隊,耗時僅 3 天打造科研級望遠鏡仿真系統、成本不足千元
阿里巴巴旗下千問辦公宣布,其產品已協助國家天文臺科研團隊搭建出一套科研級望遠鏡仿真系統。整個系統的開發僅耗時 3 天,成本不足千元。 據介紹,該仿真系統可支撐科研團隊在計算機環境中對望遠鏡的相關參數與運行狀態進行模擬,滿足科研級的模擬需求。而在此之前,同類系統通常需要委託外部軟件供應商定制開發,週期約三個月,開發成本高達數萬元。