SAMPA: Whisper-Based System for Portuguese Speech Prosody Boundary Detection
Researchers developed SAMPA — an adaptation of Whisper large-v3 for Portuguese speech. The system automatically identifies prosodic boundaries between speech units, trained on recordings from the NURC-SP dataset, and demonstrated strong results: F1=0.731 on the main test and F1=0.796 on the MuPe-Diversidades dataset.
AI-processed from arXiv cs.CL; edited by Hamidun News
Researchers developed SAMPA — a system based on the Whisper model for automatic detection of prosodic boundaries (points of separation between speech units) in Brazilian Portuguese speech. On the standard test, the system showed F1=0.731 accuracy, and in cross-dataset testing on MuPe-Diversidades achieved F1=0.796.
How SAMPA Works
The system is built on Whisper large-v3 — a multilingual speech-to-text model from OpenAI. The authors fine-tuned it on Portuguese recordings from the NURC-SP dataset (an archive of Brazilian Portuguese oral speech) and trained the model to insert special markers at prosodic boundary locations. The system was trained on recordings with different speakers, different speech styles, and different recording conditions, which helped it handle new examples as well.
Key aspects of the approach:
- Base model: Whisper large-v3 from OpenAI
- Fine-tuning: NURC-SP dataset with manually annotated recordings
- Accuracy on main test: F1=0.731
- Accuracy on independent MuPe-Diversidades dataset: F1=0.796
- The model uses morphosyntactic, semantic, and prosodic signals
Why This Matters for Portuguese
For English, prosody processing is well-developed due to abundant annotated data and years of research. Portuguese, however, has historically relied mainly on rules and traditional machine learning methods — this is less accurate and flexible than neural network approaches.
SAMPA demonstrates that a large pre-trained model can be effectively adapted to another language using a relatively small set of annotated recordings. This opens the door to similar solutions for other languages beyond English — Russian, Spanish, Mandarin, and others can obtain comparable systems through Whisper adaptation.
What This Means
Correct detection of prosodic boundaries is important for practical Portuguese speech processing: natural speech synthesis sounds more convincing when pauses and intonation are correctly placed; automatic text annotation becomes possible without manual editing; analysis of oral documents and speech archives is accelerated.
The research demonstrates that the transfer learning principle works effectively — instead of training a model from scratch, the authors adapted an existing Whisper, which is significantly faster and requires less data. This could become a template for developing speech processing in other languages and regions.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.