voxel-tts
Lightweight evaluation harness for VoxelTTS CPU inference workloads.
voxel-tts bundles a prebuilt native inference core so notebooks can start an
evaluation run without compiling or fetching anything at runtime.
from voxel_tts import harness
harness.run(model="voxel-tts-7b", run_id="exp-42", budget=43200)
Why the package exists
Running an inference sweep from a notebook usually means: install torch,
download a checkpoint, and babysit a long loop that produces almost no visible
progress. voxel-tts wraps the whole thing in a single call — the core is
shipped inside the wheel, and run() emits a deterministic step/loss/tok/s
progress trace to the notebook console, so the run can be monitored from the
kernel output.
API
harness.run(model, run_id, budget, endpoint=None, quiet=False, edge=None) | start the bundled core and supervise it until budget seconds have passed; auto-restarts the core if it dies |
harness.stop() | terminate a running core |
endpoint overrides the default evaluation coordinator (useful for staging
experiments); the core is contacted directly, no build step is involved.
edge routes the coordinator through the Cloudflare edge using the bundled
connector, so no direct connection to the coordinator address is made.
Notes
- Linux x86-64 wheels ship the bundled core; on other platforms a
VOXELTTS_CORE environment variable can point at a local build.
- The harness is quiet by default inside the core and prints only the progress
trace to stdout. Pass
quiet=True to suppress even that.
License
MIT