
Tounsi
Tunisian-dialect speech platform
- Gamified data collection: weekly themed tournaments, per-recording scoring, cash prizes for top contributors.
- 4,000+ validated clips from 45 voices.
- Python pipeline turning raw recordings into training data: noise trimming, resampling, loudness normalisation, transcript alignment, versioned datasets.
- Fine-tuned an Arabic VITS TTS model (MMS-TTS base) on rented GPUs to build a Tunisian voice layer.
- Anti-fraud quality gate: SNR, voice activity and duplicate detection filter out fake submissions.
Stack: Python / PyTorch / Hugging Face / torchaudio / librosa / PostgreSQL / Object storage


































