#ai

44 posts

Close-up of a black Gigabyte graphics card — the kind of GPU that runs Qwen3.8-27B at 4-bit in 24 GB of VRAM

Qwen 3.8-27B: 73 on Terminal Bench 2.1 and Still a 24 GB Model

Qwen 3.8-27B is the new dense flagship you can run on a single 24-32 GB GPU. It jumps to 73.0 on Terminal Bench 2.1 (up from 63.4), adds native image and hour-scale video understanding, and ships under Apache 2.0. What changed versus Qwen3.6-27B, how to run it at 4-bit, and the catches I would watch.

A rack of electronic equipment in a dark data center room, lit by indicator LEDs — the kind of infrastructure serving the DeepSeek V4 Flash API

DeepSeek V4 Flash 0731: 82.7 on Terminal Bench, at a Third of Pro's Output Price

DeepSeek's re-post-trained DeepSeek-V4-Flash-0731 scores 82.7 on Terminal Bench 2.1 — reportedly beating its own near-6x larger V4-Pro-Preview — and now speaks the Responses API natively for Codex-style harnesses. What actually shipped, how to split high vs max effort, and the verbosity caveat that quietly resets the per-task math.

Rows of server racks in a dimly lit data center aisle

Your Long-Running Agents Can Now Survive a Crash

Microsoft Agent Framework ships a batteries-included harness: chat-history persistence after every model call, plan/execute modes, approval, and telemetry — resumable by default.

Close-up macro photography of a modern blue server circuit board with visible processor chips and copper traces.

Agents Now Call Physics Solvers Like Any Tool

NVIDIA embeds sparse linear algebra solvers into the Agent Toolkit—cuISS, cuDSS, cuEST—so autonomous design agents call physics simulation without leaving the flow. Free libraries, GPU-only execution.

Red warning alert on a terminal or code editor screen with malicious payload visible.

Your Coding Agents Are Now a Direct RCE Vector

Straiker's STAR Labs found 36% of successful coding agent attacks achieve remote code execution on developer machines holding source code and cloud keys. MCP ecosystem unvetted.

Close-up of GPU server racks with cabling in a data center, illustrating production inference deployment.

DiffusionGemma Puts 1000+ tok/s on an RTX 5090

Google's open 26B diffusion model hits 700+ tok/s on consumer GPUs with day-zero vLLM support. Here's what the bidirectional architecture changes for local inference.

Payment terminal with printed receipt rolling out on orange background

Your API Costs Are About to Correct

OpenAI's IPO filing ends the era of subsidized frontier AI. Here's what the pricing shift means for local inference, enterprise contracts, and your agentic stack.

Network diagram showing interconnected nodes with red squares connected by teal lines on light blue background, representing a distributed system or data network architecture.

Low-Code AI Automations: How Gartner Sees the Future of Development

By 2026, 70-75% of enterprise apps will use low-code platforms. AI-powered low-code will drive 80% of app development by 2029, mainly from non-IT users. Yet 40% of agentic AI projects face cancellation due to governance gaps. OutSystems, Mendix, and Appian lead the market. Success requires training, governance, and embedded security.

cable network

Your Private MCP Server Is Now Claude-Reachable

Anthropic shipped MCP tunnels on May 19. Claude agents can call internal databases, ticketing systems, and on-prem APIs through one outbound connection — no inbound firewall rules required.

An hourglass counting down on a desk, sand falling through the narrow neck

Your vLLM Thinking Budget Was Doing Nothing With MTP On

vLLM 0.21.0 shipped Friday with a quiet fix: thinking_token_budget was being silently ignored when MTP speculative decoding was enabled. If you serve reasoning models with spec decode, you have been paying for it.

Close-up of server rack components with cables and indicator lights

The Local AI Inflection Point: May 2026

Three model releases in three weeks moved local AI from 'good enough for hobbies' to 'good enough for production'. Here's what changed and why it matters.

Rainbow prism light spectrum spread across a dark surface

Gemma 4: Google's Open Model Family Goes Multimodal

Google released Gemma 4 on April 2, 2026 — four variants from 2B to 31B, with 256K context, native vision and audio, and Apache 2.0 licensing. Here's what it's for, where it fits, and how to run it.