Ling 3.0 Flash is now available on AI Gateway
Ling 3.0 Flash from Ant Group is now available on AI Gateway.The model is free to use for the next three weeks, through August 3rd.Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and non-thinking modes.Ling 3.0 Flash is built for token-efficient agentic inference at production scale, doing more work within tighter token, latency, and cost budgets across multi-step agent runs. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more.AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests.Try Ling 3.0 Flash in the model playground.
- ▪Ling 3.0 Flash from Ant Group is now available on AI Gateway.The model is free to use for the next three weeks, through August 3rd.Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token.
- ▪It has a 256K token context window and runs in thinking and non-thinking modes.Ling 3.0 Flash is built for token-efficient agentic inference at production scale, doing more work within tighter token, latency, and cost budgets across multi-s
- ▪It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more.AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your O
Vercel Blog files mainly under programming. We currently carry 10 of its stories.
Opening excerpt (first ~120 words) tap to expand
Ling 3.0 Flash from Ant Group is now available on AI Gateway.The model is free to use for the next three weeks, through August 3rd.Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and non-thinking modes.Ling 3.0 Flash is built for token-efficient agentic inference at production scale, doing more work within tighter token, latency, and cost budgets across multi-step agent runs.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Vercel News.