Vāgdhenu (वाग्धेनु) — Fast WASM SIMD (MatMulInteger QUInt8) Sanskrit TTS Weights (sanskrit-tts-wasm)

🌐 Live Web App: https://h3manth.com/ai/sanskrit-tts/ 🤗 Interactive Hugging Face Space: https://huggingface.co/spaces/gnumanth/sanskrit-tts-web 🖥️ WebGPU / ONNX Weights (sanskrit-tts-onnx): https://huggingface.co/gnumanth/sanskrit-tts-onnx 📦 Source Code & npm Module: https://github.com/hemanth/sanskrit-tts-web

Native WASM SIMD MatMulInteger (QUInt8) quantized weights for Vāgdhenu, a 100% client-side, vṛtta-aware Sanskrit śloka-to-chant neural synthesis engine.

Why sanskrit-tts-wasm is the Default Browser Backend

  • 3.24× Faster ODE Step Latency (257 ms vs 835 ms per step on CPU/WASM SIMD via hardware i8x16 unsigned integer dot-product kernels).
  • 7.6× Faster Session Initialization (271 ms vs 2,066 ms) with 9× lower peak RAM (~280 MB vs 2.6 GB), eliminating CPU ConstantFolding weight expansion on mobile and desktop browsers.

Model Artifacts

File Size Description
vagdhenu_cond_q8.onnx 19.8 MB WASM INT8 Static Conditioner — QUInt8 MatMulInteger 4-layer ConvNeXt V2 text encoder + RoPE + Rank-32 static bias projection.
vagdhenu_dit_step_q8.onnx 199.1 MB WASM INT8 22-Block Flow-Matching DiT Step — DynamicQuantizeLinear + MatMulInteger (QUInt8) Diffusion Transformer with FastGelu and Rank-32 SVD AdaLN.
vagdhenu_vocos_q8.onnx 55.6 MB Fourier-Hann Vocos Vocoder — ConvNeXt backbone + exact FP32 log-magnitude/phase head + 1026-bin Fourier-Hann ConvTranspose1d synthesis (24 kHz).
baked_bank.bin + baked_bank.json 2.84 MB Pre-Baked Prosody Reference Bank — IEEE-754 FP16 mel templates across 18 classical Sanskrit vṛttas, 3 alliteration primes, and the 1,039-token phoneme vocabulary.

Credits

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support