Kimi K2.7 Code is now generally available as an integration within GitHub Copilot, enhancing code completion and development assistance capabilities for users. This expansion brings advanced coding features to Copilot's existing user base.
TL;DR
Model Releases
Tools & Products
Research Papers
Model Releases
CursorBench 3.1 is a new benchmark update providing enhanced evaluation metrics for assessing AI code editor performance and capabilities. The latest version offers improved testing standards for measuring cursor-based AI coding tools.
Claude Fable 5 offers promotional access to users interested in exploring the latest version of the AI model. This represents an opportunity for early access to new capabilities and features.
Tools & Products
ZCode serves as a harness for GLM-5.2, facilitating integration and execution of the large language model within development environments. This tool enables developers to more easily work with GLM-5.2's capabilities.
Context.dev is the web context API for AI products and agents. Scrape any URL, crawl sites, turn pages into LLM-ready Markdown, extract structured data into your own schema, capture screenshots, and retrieve logos, colors, fonts, styleguides, company data, and transaction enrichment through one API. YC-backed, no card required, and built so developers or coding agents can integrate in minutes.
Most sales AI waits for you to ask. Needle is proactive. It works like a GTM engineer on your team: it watches your pipeline and acts before you do. Spots stalled deals and drafts the follow-up, preps you before calls, keeps your CRM tidy, surfaces real buying signals. It lives in Slack and Teams, wired into HubSpot, Gmail and Gong. Unlike horizontal agents, it is built for revenue teams, acts through your permissions, and your context and memory stay portable. No lock-in. Not another dashboard.
Macro is the all-in-one workspace that combines email, messages, docs, tasks, code, agents, calls, and CRM. With team-level memory, you can query your entire workspace and never lose context.
Solaris is an AI-native transformation platform that helps organisations build AI fluency across every team. Start with a fluency test to understand where people are today, then give each team tailored learning, practical use cases, workflow challenges, champions and adoption tracking. Solaris helps companies move from scattered AI experiments to measurable capability, so AI becomes part of how work actually gets done.
Manufact, a YC S25 startup, has launched MCP Cloud, a cloud-based platform designed to streamline manufacturing operations through advanced technology integration. The service aims to provide developers and manufacturers with accessible cloud infrastructure for industrial applications.
PieterPost MCP connects AI agents to postal mail. From ChatGPT, Claude, Codex, Claude Code, or any MCP client, agents can prepare letters and postcards, use Mailbook contacts, upload attachments or postcard images, create checkout links, and track orders. It brings PieterPost online mail, API, and payment-link workflows into agent tools.
scritty is a terminal emulator that captures every CLI agent's conversation (Claude, Codex, Copilot, Antigravity, Ollama), indexes it into one searchable corpus you control, and serves it back to your agents over MCP and to you over the CLI. One session across desktop, browser, and mobile. Your captures stay on your machine.
Basedash answers questions about your data. Now it acts on them. Ask the agent to extend a trial, fix a record, or seed a demo org — it writes the SQL and runs it against any database an admin has enabled for edits. Ask it to update a Stripe subscription or create a HubSpot lead and it acts through any MCP tool you've connected. Every consequential action pauses for your approval, and every tool has its own permission. Skills chain it all into workflows. From answers to actions.
Macuse is a native macOS app that connects Claude, Codex, Cursor, Raycast, and any MCP-compatible AI client to your Mac apps. It gives AI assistants local access to Calendar, Mail, Notes, Reminders, Messages, and real app control through Computer Use.
A new CLI tool leverages embedding models to detect non-exact code duplication, helping developers identify similar code patterns that exact matching would miss. The open-source tool provides an efficient way to find and manage code reuse across projects.
Research Papers
Senior SWE-Bench is an open-source benchmark that evaluates AI agents on their ability to perform tasks at a senior software engineer level. The benchmark provides a standardized framework for assessing advanced coding abilities and complex problem-solving skills.
Research demonstrates that a single transformer layer can match the performance of full-parameter reinforcement learning training, suggesting potential efficiency gains in neural network design. This finding challenges assumptions about the depth required for effective deep learning models.
Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Things (IIoT) networks due to their suitability for resource-constrained edge deployment. Most reported results evaluate these models only within their training network, leaving behavior on unseen networks unverified. This study trains four lightweight architectures on one IIoT dataset and evaluates them, without retraining, on two structurally distinct IIoT datasets using a feature re...
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization ...
Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation. However, current systems operate within constrained search spaces or require predefined research questions, limiting their capacity for true open-ended inquiry. Furthermore, while they generate hypotheses iteratively, they largely lack the ability to explicitly synthesize their own accumulated findings to uncover complex, interconnected phenomena. ...
World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enough for online use and expressive enough to represent uncertain futures. Diffusion models offer a natural mechanism for modeling uncertain dynamics, yet their iterative inference procedure makes them difficult to use for low-latency latent planning. We bridge this gap with Value Diffusion World Models (Valdi), combining end-to-end online training for MPC with a latent diffusion dyna...
As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video reasoning) and subsequently perform fine-grained, query-conditioned tasks (intra-video reasoning) within the retrieved content, such as temporal grounding. However, existing approaches typically treat retrieval as a preprocessing step, and consequently, when the initial retrieval fails, there is no mechanism to refine...
Open-source libraries and tools are widely reused, but compatibility maintenance is expensive. Once maintainers leave, useful repositories can stop working as runtimes and dependencies evolve. We study whether LLM agents can adapt old repositories to modern environments, a task we call compatibility rescue. Unlike bug repair, compatibility rescue starts from a repository that worked in its original environment but fails after ecosystem drift. RepoRescue gives agents only the repository and its f...
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized baselines and official reference patches. Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission. We audit these i...
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result,...
Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environment mapping and the need to generate intermediate point-cloud representations. To alleviate this issue, a recent work eschews point clouds in favor of a lightweight Gaussian distribution for each object. This approximation drastically speeds up inference and enables real-time 3D scene graph generation. However, the representation has two key weaknesses. 1) Each object is approxim...
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent without assuming any knowledge of the specific executing too...
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the state-prediction separation hypothesis: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses two computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers bet...
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address thi...
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlo...
Industry News
Japan's top court has ruled that artificial intelligence cannot be listed as an inventor on patent applications, establishing important legal boundaries for AI intellectual property rights. This decision clarifies that patents must have human inventors in Japan's legal framework.
Spain has ordered a blacklist of Palantir, prohibiting the controversial data analytics company from working with both public and private organizations in the country. This action reflects growing concerns about data privacy and surveillance in European markets.
Meta has implemented spending caps on internal AI token usage to manage computational costs and resource allocation. This decision reflects efforts to optimize AI infrastructure spending across the company.
OpenAI is reportedly in early discussions with the US government about potentially granting a 5% equity stake, indicating possible governmental interest in maintaining influence over the influential AI company. The talks suggest evolving discussions around AI governance and public-private partnerships.