bosonai-higgs-audio-STT-3
BosonAI · Operations & Productivity
Certification per Microsoft Marketplace.
Evidence tier Source Confirmed · 8 captures on record
What the publisher says
As described on Microsoft Marketplace.
Test plan for boson AI Higgs audio STT (ASR) model v3. bosonai-higgs-audio-STT-3 is the latest audio understanding model from Boson AI, succeeding [Higgs Audio v1](https://www.boson.ai/blog/higgs-audio), higgs-audio-STT v3 marks a return to understanding with state-of-the-art Speech-to-Text (STT) capabilities. This model utilizes a LLM backbone with a specialized encoder trained on top of a whisper-form-encoder. It delivers industry-leading performance in Automatic Speech Recognition (ASR) and Speech Translation (AST), outperforming `whisper-v3-large` on key languages like English, Spanish, and Chinese by a large margin. Higgs Audio v3 STT introduces significant architectural and data-centric advancements over previous generations. We implemented architectural changes to support **chunk-prefill**, allowing users to process audio every 4 seconds or based on Voice Activity Detection (VAD) chunks. This drastically cuts response latency, making it ideal for real-time applications compared to traditional full-sequence processing. The model supports ASR both with and without language hints. In "no hint" mode, it dynamically adapts to the input language, enabling seamless transcription of multilingual audio without prior configuration—a critical feature for production environments. Through advanced data augmentation techniques, Higgs Audio v3 STT demonstrates exceptional robustness in challenging acoustic environments, maintaining high accuracy even with background noise or poor recording quality. The typical use cases for bosonai-higgs-audio-STT-3 are 1) High-Accuracy Transcription: which converting speech to text for meetings, lectures, and interviews with performance exceeding current state-of-the-art models; 2) Real-time captioning. Leveraging the streaming/chunk-pre-fill capabilities to provide low-latency captions for live broadcasts or streaming. 3) Multilingual translation: performing automatic speech translation (AST) to directly translate spoken content to other languages.; 4) Dynamic language processing: transcribing audio streams with unknow or mixed languages using model's dynamics adaptation capabilities.
Preview
1 imageAgent build and provenance
Sign in to see the provenance.
The evidence, the layer-by-layer tracing, the risk basis, and the cross-marketplace links are open to signed-in accounts.
Sign inCompliance
- FedRAMPConfirmedNot listed90%, registry-checkedNo FedRAMP Marketplace entry matched this vendor's domain, checked 2026-08-27registry recordas observed 2026-08-27
Confirmed means matched to a public authoritative registry. Claimed means the vendor or its listing states it, not yet cross-checked. A framework not shown was not found in any source we hold, which is not evidence against it. Not listed means a scoped registry check found no match for this vendor's domain: a No is a scoped registry check, not a compliance judgment. Confidence bands: 95% domain-verified, 90% registry-checked, 80% self-attested, 70% weak signal. Self-attested items marked “vendor's site” are gathered from the vendor's own website and are not verified by us.
Vendor
External enrichment
Plans and pricing as listed
1 listed- paygo-surcharge-a10-gpu : $0.36 per gpu hour
- paygo-surcharge-h100-gpu: $0.36 per gpu hour
- paygo-surcharge-a100-gpu: $0.36 per gpu hour
Sources
Publisher resources
1 linkEvidence risk is the share of the build you cannot see before you deploy, not a security rating. Sign in to see the layer-by-layer basis for this band.


