This session is a technical walkthrough of what it takes to run Open Frontier AI in production: inference, SFT, custom speculative decoding, and dedicated endpoints, on Nebius Token Factory.
What to expect
A technical presentation and demo of how to run full-stack AI Engineering Pipeline with Nebius Token Factory.
1. Choosing your model stack — open weights vs. proprietary APIs
A decision framework, not a vendor pitch. When open-weight models win on cost, latency, data control, or flexibility, and when they don't. What Kimi K3 changes about that calculus now that a 2.8T MoE model ships at open weights and scores 57 on Artificial Analysis' Intelligence Index, just two points behind GPT-5.6 Sol.
2. Building on Kimi K3
The architecture that makes it work: 2.8T total parameters activating only 16 of 896 experts per token, Kimi Delta Attention, native vision, a full 1M-token window. Where K3 outperforms alternatives on long-horizon coding and agentic tasks, where it doesn't, and integration patterns through Token Factory's OpenAI-compatible API.
3. Full-stack AI engineering — demoed across the whole pipeline
4. Open floor
Talk to us and your peers about the hardest AI Problems you are trying to solve.
Who this is for
CTOs, VPs of Engineering, AI Architects, AI FDEs and senior AI engineers at growth-stage startups building AI-native or AI-augmented products — typically Series A and beyond.