Speech research

Haryanvi TTS

A reproducible data and training workflow exploring natural Bangru/Haryanvi speech synthesis with IndicF5 on accessible GPU infrastructure.

Project state
Research in progress
Documented
2026

Dataset and training workflow validated in bounded experiments; production readiness is not claimed

Bangru / Haryanvi voice lab

Sample review · 03

00:00Full precision sample00:08

01

Audit corpus

02

Gate training

03

Listen & review

Research artifact · production readiness is not claimed

Product thesis

Regional-language voice quality starts with disciplined data, not a bigger demo.

The constraint

Start with the real friction.

Haryanvi and its Bangru speech variety have limited high-quality text-to-speech infrastructure, while community recordings often arrive with inconsistent metadata, duplicate content, and unlinked audio.

The response

The project builds a gated workflow around dataset reconciliation, deterministic splits, pretrained-weight validation, small overfit tests, and persisted Colab artifacts before any larger training claim is made.

System map

How the pieces fit.

  1. 01

    Audit

    Reconcile metadata-linked recordings, normalize transcripts, isolate rejected and orphan files, and generate review ledgers.

  2. 02

    Gate

    Validate model access, checkpoint structure, parameter coverage, dataset loading, and a tiny full-precision training run.

  3. 03

    Train and review

    Run bounded Colab experiments, persist artifacts to Drive, and listen to generated samples before considering a larger run.

What it does

Concrete capability.

  • Deterministic audit of transcripts, recordings, duplicates, and orphan files
  • Reproducible free-GPU training path based on a pinned IndicF5 source
  • Checkpoint conversion and pretrained-parameter coverage checks
  • Baseline audio, smoke training, overfit gates, and persisted review reports

Product principles

The choices behind it.

  1. 01Treat the metadata ledger as the authority before training
  2. 02Quarantine unmatched audio instead of silently adding it to the corpus
  3. 03Prove a tiny run can learn before spending time on scale
  4. 04Preserve reports, samples, and checkpoints so results can be inspected

Evidence, not theatre

What can be said clearly.

2,768
metadata-linked recordings audited
15/15
observed local workflow checks
T4
accessible Colab training target

Built with

  • Python
  • PyTorch
  • IndicF5
  • Hugging Face
  • Google Colab
  • FFmpeg
  • Dataset audit tooling

Continue the conversation

Interested in the product—or the system behind it?

Discuss regional-language voice AI

Next project

KissPDF