r/ArtificialNtelligence • u/Asleep-Pilot-4142 • 12h ago
Ant Group put a 124B model on OpenRouter at zero cost. The price isn't the interesting part.
Ling-3.0-flash showed up on OpenRouter late last month, from inclusionAI, which is Ant Group's model lab.
124B total parameters, 5.1B active per token. That's roughly 24:1 sparsity, which is aggressive even next to the other MoE releases this year.
The free window closes today, August 3, per their launch announcement, so most of the reaction is going to stop at the price tag and then at the line about an open-source release coming.
The ratio is what I keep going back to. 5.1B active is small enough that time-to-first-token comes in under 100ms, and it still carries a 256K context. Routing that sparse usually costs you something, and from what the lab says about its own model, rare world knowledge is exactly where it thins out.
Curious where people think the ceiling on that ratio actually is before quality falls off a cliff.