Bosonai-higgs-audio-3-instruct
BosonAI · Operations & Productivity
Certification per Microsoft Marketplace.
Evidence tier Source Confirmed · 9 captures on record
What the publisher says
As described on Microsoft Marketplace.
Higgs Audio v3 Instruct is Boson AI’s production-quality,
audio-native instruct model (aka Audio LLM): a 14B instruction-tuned LLM that
Show the rest of the publisher’s description (30 more lines)
can understand audio, text, or both, and generate high-quality,
instruction-following text responses. It functions as a strong text LLM when
given text alone, while bringing native audio understanding to speech and
multimodal audio-text inputs.
Unlike today’s omni-chat audio LLMs — such as GPT-4o
audio mode, Gemini 2.5 audio, and Qwen2.5-Omni — Higgs Audio v3 Instruct is
fine-tuned specifically for **voice-agent reflexes**: audio-native tool
calling, multi-turn state tracking, and interruption-aware instruction
following. These capabilities are trained directly into the model weights,
rather than prompt-engineered on top.
Boson AI Higgs Audio Instruct audio-in model is now
competitive with strong text models on instruction following, unlocks next
level of intelligence efficiency for voice agents that previously had to choose
between audio understanding and Instruction-Following. Higgs Audio v3 Instruct
Audio-in LLM is capable of Text-Model-Grade Instruction Following. Scored 85.5
on IFEval bench, scored 30.4 on IFBench bench, scored 27.6 on MultiChallenge
bench, and scored 31.3 on MultiChallenge-Audio bench.
Higgs-Audio-Instruct model is capable of audio-native
function calling. Function/tool calling is now part of the model's behavior,
not a bolted-on prompt hack. The API supports standard OpenAI-style function
calling. The release checkpoint scored 22.1 success / 55.2 call_acc on the
audio-converted ComplexFuncBench — the strongest tool-use surface we have
shipped on a 14B audio LLM.
The model holds the instruction frame across turns and
interruptions and stays in-flow rather than resetting after every turn. Scored
27.9 on AudioMultiChallenge bench, Scored 46.1 / 69.9 / 88.3 on Interruption-v2
(true-follow / re-query / resume).
Chunk-prefill audio input — VAD-segmented at up to
4-second chunks, 16 kHz, robust to noise and accent. ASR / AST is supported as
an inherited capability.
Preview
1 imageAgent build and provenance
Sign in to see the provenance.
The evidence, the layer-by-layer tracing, the risk basis, and the cross-marketplace links are open to signed-in accounts.
Sign inCompliance
- FedRAMPConfirmedNot listed90%, registry-checkedNo FedRAMP Marketplace entry matched this vendor's domain, checked 2026-08-27registry recordas observed 2026-08-27
Confirmed means matched to a public authoritative registry. Claimed means the vendor or its listing states it, not yet cross-checked. A framework not shown was not found in any source we hold, which is not evidence against it. Not listed means a scoped registry check found no match for this vendor's domain: a No is a scoped registry check, not a compliance judgment. Confidence bands: 95% domain-verified, 90% registry-checked, 80% self-attested, 70% weak signal. Self-attested items marked “vendor's site” are gathered from the vendor's own website and are not verified by us.
Vendor
External enrichment
Plans and pricing as listed
1 listed- paygo-surcharge-a100-gpu: $5.40 per gpu hour
- paygo-surcharge-h100-gpu: $5.40 per gpu hour
- paygo-surcharge-a10-gpu : $5.40 per gpu hour
Sources
Publisher resources
1 linkEvidence risk is the share of the build you cannot see before you deploy, not a security rating. Sign in to see the layer-by-layer basis for this band.


