The constraint
Start with the real friction.
Haryanvi and its Bangru speech variety have limited high-quality text-to-speech infrastructure, while community recordings often arrive with inconsistent metadata, duplicate content, and unlinked audio.
Speech research
A reproducible data and training workflow exploring natural Bangru/Haryanvi speech synthesis with IndicF5 on accessible GPU infrastructure.
Dataset and training workflow validated in bounded experiments; production readiness is not claimed
Bangru / Haryanvi voice lab
Sample review · 03
01
Audit corpus
02
Gate training
03
Listen & review
Research artifact · production readiness is not claimed
Product thesis
“Regional-language voice quality starts with disciplined data, not a bigger demo.”
The constraint
Haryanvi and its Bangru speech variety have limited high-quality text-to-speech infrastructure, while community recordings often arrive with inconsistent metadata, duplicate content, and unlinked audio.
The response
The project builds a gated workflow around dataset reconciliation, deterministic splits, pretrained-weight validation, small overfit tests, and persisted Colab artifacts before any larger training claim is made.
System map
01
Reconcile metadata-linked recordings, normalize transcripts, isolate rejected and orphan files, and generate review ledgers.
02
Validate model access, checkpoint structure, parameter coverage, dataset loading, and a tiny full-precision training run.
03
Run bounded Colab experiments, persist artifacts to Drive, and listen to generated samples before considering a larger run.
What it does
Product principles
Evidence, not theatre
Built with
Continue the conversation
Next project
KissPDF