Inflect-Micro-v2: complete voice in 9.36M parameters
Posted by nateb2022 1 day ago
Comments
Comment by modinfo 1 day ago
here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!
Comment by yjftsjthsd-h 1 day ago
> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
Comment by semiquaver 1 day ago
Comment by yjftsjthsd-h 1 day ago
Comment by NetOpWibby 1 day ago
Comment by tmaly 1 day ago
Comment by fastball 1 day ago
Comment by billdueber 1 day ago
Comment by SamPatt 19 hours ago
I learned about all of these projects on HN at one point or another.
Comment by eightysixfour 1 day ago
Comment by _davide_ 1 day ago
Comment by K0balt 1 day ago
Comment by sudb 1 day ago
Comment by da-x 1 day ago
Comment by StilesCrisis 1 day ago
Comment by g58892881 1 day ago
Comment by StilesCrisis 1 day ago
On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?
Comment by g58892881 1 day ago
Comment by jsomedon 1 day ago
Comment by itake 1 day ago
IMHO, its at about the same quality level of historic TTS tools.
Comment by stavros 1 day ago
Comment by phoenixranger 1 day ago
Comment by mcbetz 1 day ago
Comment by fintuner 1 day ago
Comment by amelius 1 day ago
Comment by zenith605 1 day ago
Comment by afdsaifdoi 1 day ago