Last week in AI was just craazzyyy.
1️⃣ Google’s Gemini 3 + Antigravity IDE Google didn’t just ship a strong new frontier model. They’re going after the whole developer ecosystem with Antigravity, a new VS Code fork optimised for agentic workflows.
2 min read
Originally posted on LinkedIn · 7 reactions · 0 comments · View original →
Last week in AI was just craazzyyy. Four product launches really stood out 👇
1️⃣ Google’s Gemini 3 + Antigravity IDE
Google didn’t just ship a strong new frontier model. They’re going after the whole developer ecosystem with Antigravity, a new VS Code fork optimised for agentic workflows.
If developers do adopt this (a big if, since engineers are very loyal to their IDEs), the competitive question shifts from “whose model is best?” to “which environment is the default place where software gets built?”
Either way, I love the competition at both layers:
- Frontier coding models: Sonnet 4.5, GPT‑5.1‑Codex‑Max, Gemini 3 Pro
- Orchestration layer: Claude Code, Copilot, Cursor, Windsurf, Antigravity, Cline, RooCode, OpenCode, Augment, Warp, etc.
At this stage, I wouldn’t be rushing into a hard vendor lock-in.
2️⃣ Google’s Nano Banana Pro: visual reasoning model, not just pretty pictures
Nano Banana Pro is a visual reasoning model that can render accurate UI layouts: headings, labels, menus, structures, and multilingual content.
This is where image generation starts to move from marketing assets to core product and UX workflows: landing pages, onboarding flows, dashboards, internal tools.
I wonder whether this would accelerate the feedback loop between design and product teams? Also, this will put significant pressure on OpenAI and Anthropic to advance their multimodality.
3️⃣ Meta’s SAM 3: natural language control of vision
SAM 3 can segment and track objects in images/video using natural language prompts: “find everyone wearing yellow”, “identify who is not using a safety vest”, “track all forklifts in this feed”.
Vision models are shifting from pixel geometry to semantic perception. That’s a big deal. Vision streams basically become searchable data.
4️⃣ Marble: the first frontier multimodal world model
World Labs released a 3D tool that can generate stable, editable worlds suitable for games, film previs, simulations and robotics. 🤯
Imagine a virtual world with infinite, dynamic rendering happening on the fly, as you navigate it.
I’m still trying to make sense of what that means for enterprise software… but for humanoid robots, this compresses years of data collection that would need to happen in the wild (e.g., Tesla self-driving) and accelerates deployment into the real world.