🤔
Alex Veremeyenko@alex_verem·Albert Gu aka the Mamba guy has the top ranked voice model now. How much of that is the architecture and how much is they just trained it better? x.com/krandiash/stat…
chat with your model with a single identity across languages and accents!
Karan Goel@krandiash·Introducing Multilingual Voices We partnered with Baklavastory, our neighbor in the Mission, to show what it sounds like when a business's character and warmth stay consistent, no matter who calls or what language they speak. It's easy to translate speech, but much harder to preserve a single identity across languages. Every language has different rhythms, tones, and emphasis. When you change the language, identity tends to get lost in that shift, and your voice ends up sounding like someone else. With Multilingual Voices, pick a voice and keep one brand identity in every market you serve: t.co/CDgBhsRR2o
poor ronald 😭
Cartesia@cartesia·We regret to inform our AI researcher Ronald that AI has replaced him. Sorry Ronald. Here's our team trying to tell his real voice from his cloned one. They were not good at it.
misleading paper title, and even the phrasing in the tweet is still a misnomer. this is not about "post-training" but "retrofitting" (cross-architecture distillation) fwiw i've been bearish on distilling Transformers to recurrent models for a while, they're too different; architectural innovations just need to be trained from scratch
Alexia Jolicoeur-Martineau@jm_alexia·Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: arxiv.org/abs/2608.28444
sonic-3.6 is out, with a large improvement over (the already #1) sonic-3.5 in just a few months! the research team's focus on fundamentals is accelerating progress at the frontier of architectures and audio
Cartesia@cartesia·Sonic-3.6 is now generally available. In January we made a bet: stop tuning the existing paradigm, rebuild from the architecture up. How we topped our own best model in two months → cartesia.ai/blog/sonic-3.6…
the team continues cooking 👩🍳 this is now an unprecedented gap on both the super competitive Provider Voices leaderboard as well as the newer Controlled Voices leaderboard (a stronger benchmark that can't be benchmaxxed, requiring truly better algorithms) Cartesia's TTS model is not only the best, it continues improving at a faster rate than the competition - Sonic 3.5 was released only 3 months ago. ig Sonic now means both the fastest model and the fastest team 💀
Artificial Analysis@ArtificialAnlys·Cartesia's Sonic 3.6 takes the #1 spot on both the Provider Voice and Controlled Voice Artificial Analysis Speech Arena leaderboards, surpassing Speechify AI's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus, with Sonic 3.5 holding #2 on Controlled Voice Sonic 3.6 is the latest TTS model from @cartesia, supporting 40+ languages including English, Hindi, Spanish, French, German and Japanese. In the TTS Arena, the model has been preferred by voters and has demonstrated natural speech and adaptive pacing across conversations. Key takeaways: ➤ Controlled Voice: Sonic 3.6 takes #1 on the Controlled Voice Arena with an Elo of 1,144 (+18/-18) across 1,339 appearances, ahead of Cartesia's own Sonic 3.5 at 1,099 and ElevenLabs' Eleven v3 at 1,060 ➤ Provider Voice: Sonic 3.6 also takes #1 on the Provider Voice Arena with an Elo score of 1,286 (+19/-19) based on 1,300 arena appearances, placing it ahead of Simba 3.2 at 1,238 and Qwen-Audio-3.0-TTS-Plus at 1,237 ➤ Pricing: At $49 per 1M characters via the Cartesia platform, Sonic 3.6 is the most expensive of the top three models on quality, roughly 5x Speechify AI's Simba 3.2 at $10 and nearly double Alibaba's Qwen-Audio-3.0-TTS-Plus at $27.59, though still half the price of ElevenLabs' Eleven v3 at $100 ➤ Speed: Sonic 3.6 processes 136.1 characters per second of generation time, compared to 46.7 for ElevenLabs' Eleven v3 and 24.2 for Google's Gemini 3.1 Flash TTS
Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS x.com/cartesia/statu…
Cartesia is hosting an ICML party tomorrow (Thursday) night! it'll go late, but come early or i may have to bounce you again luma.com/gy1b4ryq
Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adjectives that say what a sentence is about" x.com/allen_ai/statu…
Rather than interleaving layers naively, a more fine-grained approach to hybrid models is to allow hybridization across the sequence models within a single layer. The fact that softmax attention and linear attention use similar underlying projection parameters allows switching between different mixers in a single generation, for the best of both worlds.
Congrats to Henry and Naomi - they’ve been so on top of the space and super helpful as collaborators too! x.com/HenryYin_/stat…