
@_albertgu
assistant prof @mldcmu. chief scientist @cartesia_ai. leading the ssm revolution.
Evaluations are difficult and vague for all generative models, and benchmarks only capture a small slice. Our blog post dives into the nuances for TTS x.com/cartesia/statu…
Cartesia is hosting an ICML party tomorrow (Thursday) night! it'll go late, but come early or i may have to bounce you again luma.com/gy1b4ryq
Transformers are better at copying, while RNNs are better at modeling "meaning-bearing words—the nouns, verbs, & adjectives that say what a sentence is about" x.com/allen_ai/statu…
Rather than interleaving layers naively, a more fine-grained approach to hybrid models is to allow hybridization across the sequence models within a single layer. The fact that softmax attention and linear attention use similar underlying projection parameters allows switching between different mixers in a single generation, for the best of both worlds.
Congrats to Henry and Naomi - they’ve been so on top of the space and super helpful as collaborators too! x.com/HenryYin_/stat…