Mistral's Robostral Navigate represents a breakthrough in robotics navigation, achieving state-of-the-art performance in autonomous movement and spatial understanding.
TL;DR
Model Releases
Tools & Products
Research Papers
Model Releases
GPT-5.6 Sol, alongside companion models Terra and Luna, will launch to the public this Thursday, marking a significant AI model release from OpenAI.
Grok 4.5 is the latest iteration of xAI's conversational AI model, offering enhanced capabilities for reasoning and real-time information processing.
SWE-1.7 has achieved performance levels approaching GPT 5.5 and Claude Opus, demonstrating significant advances in software engineering AI capabilities.
Migtissera released Tess-4-27B, a new AI model with 27 billion parameters designed for advanced text processing and reasoning tasks. This release represents progress in open-source large language model development with a focus on efficiency and performance.
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
Tools & Products
Chatto, an open-source conversational AI model, is now available to developers for customization and deployment on their own infrastructure.
Willow is launching two new voice AI models: Willow Frontier Pro and Willow Frontier Mini. Frontier Pro is our most powerful dictation model. It is built for people who want fast, accurate, and polished writing anywhere they work. Frontier mini is lightweight and completely free to use. On our free plan, users get unlimited access to this model. It's faster and more accurate than other AI dictation tools.
Geosql is a new skill/plugin for Claude that enables processing and querying geospatial data, expanding the model's capabilities for location-based analysis.
LemonLime lets teams automate their workflows in minutes with a single click. It connects to your existing tools, studies your business, and self-creates specialized AI agents and automations that support your team. Don’t know where to start? LemonLime helps with that, too, automatically surfacing suggested automations that you can implement with a single click.
PopTask just went universal. type a messy thought like "gym mon wed fri 6am" and it becomes a scheduled task in about 3 seconds, no pickers, no forms. on iphone + ipad you get home and lock screen widgets, a live activity + dynamic island counting down your next task, control center, and hands-free siri even in the car. on mac it lives in the menu bar (⌘⌃P). everything syncs across your devices through your own icloud, near-instant. on-device, private, 9 languages. free to start
One command-line tool to scaffold, evaluate, and deploy AI agents on Google Cloud — built to be driven by your coding agent (e.g Antigravity, Claude Code, Codex). Scaffold a production-ready project, evaluate against a real signal, and ship to Agent Runtime, Cloud Run, or GKE or anywhere else!
Universal-3.5 Pro is AssemblyAI's most accurate speech-to-text model, now available at our Realtime & Async endpoints. It transcribes every conversation exactly as it's heard—code-switching across 18 languages, our most accurate speaker diarization yet, and contextual prompting to steer results.
Stop settling for robotic dubbing. Cutrix uses an Agentic workflow to translate videos while preserving the speaker's original emotion and natural pacing. Experience hyper-natural alignment without the steep learning curve. Sign up for free credits today!
Eodly reads Slack, Telegram, Discord, GitHub and Linear, and sends founders one sourced page each evening: who shipped, who's quiet, who's slipping, and any status that doesn't match reality. Your team never logs in. A chief of staff, not surveillance.
Cloudflare Meerkat introduces a globally distributed consensus mechanism, enhancing network reliability and security across multiple geographic regions.
Research Papers
Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revision. We introduce SWE-Review, a framework for closing this loop with agentic code review. Given an issue and an AI-generated PR, a reviewer agent explores the repository, decides whether the PR should be accepted, and provides structured feedback for revision. We evaluate this setting with our propos...
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed t...
Scaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware and workspaces. We introduce digital teleoperation, a paradigm that decouples data collection from physical constraints by replacing the real robot with a generative world model. In this framework, an operator's hand-pose stream drives a robot-centric generative world model to synthesize high-fidel...
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours ...
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of ...
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy, noise, and uneven coverage. Existing data selection methods often assess demonstrations at either the trajectory or state-action level, missing the reusable structures that compose long-horizon behaviors. In this paper, we propose SIEVE, a structure-aware data selection method for VLA imitation lea...
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neur...
Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world models offer a fundamentally different paradigm. Rather than explicitly authoring every component of a virtual environment, these models autoregressively synthesize future observations conditioned on the current world state and user interactions, enabling playable worlds to be ge...
Late-interaction retrieval models that use the MaxSim similarity function have shown strong empirical performance, often outperforming single-vector dense and sparse retrieval models. Despite these empirical findings, little is known about the theoretical representation power of MaxSim and how it compares to other retrieval approaches. This paper shows by construction that MaxSim similarity can exactly replicate the inner product between any two non-negative k-sparse vectors with possibly infini...
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extens...
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significant...
We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-specu...
Tutorials
OpenAI Academy and the Walton Family Foundation are bringing hands-on AI Skills Jams to help K–12 educators build practical AI skills for the classroom.
Industry News
Researchers discovered a security vulnerability in GitHub's AI Agent that could be exploited to leak private repository information, highlighting important AI safety concerns.