Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs
▲ 19 r/OpenSourceeAI+1 crossposts

Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs

Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs

Here's what's actually in the release:

1. Three drafters, one per target model

→ LFM2.5-1.2B-Instruct, LFM2.5-2.6B, LFM2.5-8B-A1B

→ Each drafter is ~300M params (295.7M / 327.7M / 327.7M)

→ 5 attention layers, block size 9, ships no vocab weights

2. The speedups are real but uneven

→ 3.18x on H100 for 8B-A1B on MATH500 (428 → 1362 tok/s)

→ 2.87x on an M4 Max for 1.2B-Instruct on HumanEval (136 → 389 tok/s)

→ 2.67x H100 mean for 2.6B (323 → 864 tok/s)

→ Same 8B-A1B model drops to 1.29x on GSM8K, same GPU

3. Speedup tracks acceptance rate, not model size

→ 8B-A1B accepts 8.27 of 10 tokens per step on MATH500

→ It accepts 4.02 on GSM8K

→ That single number explains the 3.18x vs 1.29x gap

4. Output quality does not move

→ Under greedy decoding, a draft token is kept only if it matches the target's distribution

→ On rejection, the target's own token takes its place

→ The emitted sequence is identical to baseline by construction

> Full analysis: https://www.marktechpost.com/2026/08/20/liquid-ai-releases-lfm2-5-dspark-draft-models-that-deliver-up-to-3-18x-faster-decoding/

> LiquidAI/LFM2.5-1.2B-Instruct-DSpark: https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark

> LiquidAI/LFM2.5-2.6B-DSpark: https://huggingface.co/LiquidAI/LFM2.5-2.6B-DSpark

> LiquidAI/LFM2.5-8B-A1B-DSpark: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-DSpark

Technical details: https://www.liquid.ai/blog/lfm2.5-dspark

marktechpost.com
u/ai-lover — 17 hours ago

ByteDance Seed and Tsinghua AIR Introduces CUDA Agent: A Large-Scale Agentic RL System for CUDA Kernel Generation

ByteDance Seed and Tsinghua AIR have released CUDA Agent, an agentic reinforcement learning system that trains a large language model to write GPU kernels that beat a compiler. The gap it targets is narrow but stubborn: frontier models already produce correct CUDA, they just produce slow CUDA. On KernelBench, the base model Seed1.6 passes 74.0% of tasks yet outruns torch.compile on only 27.2% of them, at a 0.69× geometric-mean speedup which means its kernels are, on average, slower than what the compiler generates on its own. CUDA Agent closes that gap by putting the model inside a real CUDA development environment with profiling, correctness checks and a permission-locked sandbox, then training it with PPO for 150 steps at a 131,072-token context. The result is a 98.8% pass rate and a 96.8% faster-than-torch.compile rate across the 250-task benchmark, at 2.11× geomean over compile — roughly 40 points ahead of Claude Opus 4.5 and Gemini 3 Pro on the hardest Level-3 split.

Full analysis: https://www.marktechpost.com/2026/08/17/bytedance-seed-and-tsinghua-air-introduces-cuda-agent-a-large-scale-agentic-rl-system-for-cuda-kernel-generation/

Paper: https://arxiv.org/pdf/2602.24286v1

u/ai-lover — 3 days ago
▲ 9 r/AIDeveloperNews+1 crossposts

DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

Here are some key takeaways:

  1. There is no privileged core → Models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI are all Cordis plugins → Any of them can be selected, swapped, or extended in configuration, without editing harness source

  2. Four runtime modes, one kernel → Standard, Code, Minimal, Creator — each loads a different default plugin set → Minimal keeps two tools, persistent bash and str_replace_editor, for benchmarking models in a bare environment

  3. Every run is traceable → An append-only session log records system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection → Resume, fork, search, and replay all operate on the same event stream

Full analysis: https://www.marktechpost.com/2026/08/17/deepseek-ai-releases-deepseek-harness-in-developer-preview/

Repo: https://github.com/deepseek-ai/deepseek-harness

marktechpost.com
u/ai-lover — 4 days ago

Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

→ Built on 3.6 Flash with algorithmic improvements to the reasoning core. Same 1M context, 64K output, March 2026 cutoff.

→ The gains concentrate in three places: software engineering, document-heavy knowledge work, and web development. The sharper argument is price.

Performance:

→ FrontierCode 1.1: 43.6% vs 34.4%

→ DeepSWE v1.1: 65.3% vs 48.6%

→ WebDev Arena: 1588 Elo vs 1538

→ AutomationBench: 30.4% vs 17.0%

→ GDP.pdf: 34.0% vs 22.0%

Full analysis: https://www.marktechpost.com/2026/08/13/google-ai-just-released-gemini-3-7-flash/

Technical details: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/

u/ai-lover — 8 days ago
▲ 27 r/OpenSourceeAI+1 crossposts

The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model

The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model. It's an open weights world model for video, real-time apps, and physical AI — optimized with NVIDIA to run on RTX GPUs and DGX Spark.

Here's what stood out:

1. The speed numbers are the story

In LTX's published image-to-video benchmark (10-second clip):

→ 6.8 seconds on-prem (2x NVIDIA GB200)

→ 23.7 seconds via the LTX API

→ 52–70 seconds for the fastest closed rivals (Omni Flash, Grok 1.5, Veo 3.1)

→ 398 seconds for Kling 3.0 Pro — that's 58.5x slower

On-prem generation finishes faster than the clip itself plays.

2. Multishot consistency fixes the real blocker

Earlier open models generated each shot separately, so characters drifted between cuts — unusable for actual campaigns. LTX-2.5 renders the full sequence as one output, holding character, scene, and voice across cuts. A custom Gemma 4 backbone handles complex, multi-subject prompts.

3. Diffusion Fidelity Rendering is a smart cost tradeoff

→ Motion and structure built in an 8x temporally compressed latent space

→ Full detail spent only on high-fidelity keyframes

→ Keyframe count adapts to scene complexity

Quality lands where it matters without full render cost on every frame.

Full analysis: https://www.marktechpost.com/2026/08/11/the-video-production-stack-now-fits-on-one-desk-ltx-2-5-launches-as-nvidia-accelerated-open-weights-world-model/

Model weight: https://huggingface.co/Lightricks

Technical blog: https://blogs.nvidia.com/blog/local-ai-open-source-models-agents-nemotron/

u/ai-lover — 10 days ago

webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware

webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware

At just 3 billion parameters, TwiL-LM3 outperforms OpenAI’s GPT-OSS-120B on 4 of 5 formal reasoning benchmarks while running efficiently on consumer hardware. That’s 40× fewer parameters, 2.6× faster inference, and state-of-the-art performance in the reasoning tasks that power reliable tool calling, code generation, structured outputs, and AI agents.

Full analysis: https://www.marktechpost.com/2026/08/10/webai-releases-twil-lm-a-1-7b-and-3b-formal-logic-model-family-for-autoformalization-on-local-hardware/

Model weights: https://huggingface.co/webAI-Official/TwIL-LM3

Technical details: https://www.webai.com/blog/webai-releases-twil-lm-a-family-of-formal-logic-models-that-outreason-a-120b-model-and-run-on-an-iphone

u/ai-lover — 10 days ago
▲ 34 r/OpenSourceeAI+1 crossposts

Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU

Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU

Meta has released Muse Glimmer, a 30-billion-parameter multimodal model distilled from Muse Spark. It is tuned for always-on local agent workflows, and ships under Apache 2.0. A 30B model normally needs over 55 GB of memory at full precision. Meta compresses it to roughly 4-bit, then adds block-level speculative decoding so it answers fast enough to sit inside a real agent loop. The result runs on one consumer GPU or a Mac, with no network call....

Model and training

Muse Glimmer is a dense causal transformer with a dedicated perception encoder. Total parameters are roughly 30B, including the vision tower. Grouped-query attention uses 32 query heads and 2 KV heads. Attention repeats a [Local, Local, Local, Global] pattern with a 2,048 sliding window. RoPE is applied to local layers only, with theta 500,000. The vision side is a ~1.8B ViT-G/14 perception encoder accepting up to 4,096 visual tokens per image. Context length is 131,072+, vocabulary is 202,048 tokens, and the knowledge cutoff is January 4, 2026. Input is text and image; output is text. Audio is not supported, and video is processed as individual frames.

Training ran in three phases:

  • Pre-training used logit distillation on Muse Spark’s outputs.
  • Mid-training added longer-context, agent-heavy data with richer reasoning traces.
  • Post-training combined supervised fine-tuning with on-policy distillation and reinforcement learning across general, reasoning, coding, and agentic domains.

Full analysis: https://www.marktechpost.com/2026/08/10/meta-ai-releases-muse-glimmer/

Model weight: https://huggingface.co/collections/meta-models/muse-glimmer

u/ai-lover — 11 days ago
▲ 29 r/OpenSourceeAI+1 crossposts

Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

It's a policy-adaptive multimodal safety classifier. Most guardrail models bake a fixed harm taxonomy into their weights, so re-targeting one means retraining. This one takes the policy as a plain-language question at inference time.

Here's what's actually interesting:

𝗠𝗼𝗱𝗲𝗿𝗮𝘁𝗶𝗼𝗻 𝗿𝗲𝗱𝘂𝗰𝗲𝗱 𝘁𝗼 𝗼𝗻𝗲 𝘆𝗲𝘀/𝗻𝗼 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻

Three fields per request. <Instruct> sets evaluation context and strictness. <Query> states the policy as a single yes/no question. <Document> holds the content — a prompt, a response, a prompt-response pair, or an image with optional text.

At inference the model unembeds only toward the yes and no token IDs, softmax-normalizes them, and thresholds at 0.5. One forward pass, one token, continuous score.

𝗧𝗲𝘅𝘁 𝗮𝗻𝗱 𝗺𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗿𝗲𝘀𝘂𝗹𝘁𝘀

→ 84.9% average text F1 — ties GPT-OSS-Safeguard-20B

→ 83.8% multimodal F1 vs 77.6% for OmniGuard-7B

→ VLGuard 97.7, UnsafeBench 81.8, HarmBench prompt 99.4

→ 91.5% refusal detection overall

𝗔𝗱𝗮𝗽𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗯𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸

→ Shieldstral-3B: 91.3% F1

→ GPT-OSS-Safeguard-20B: 94.1%

→ Nemotron-3.5-Safety-4B: 91.8%

Full analysis: https://www.marktechpost.com/2026/08/07/mistral-ai-releases-shieldstral-1-0-3b/

Model weight: https://huggingface.co/mistralai/Shieldstral-1.0-3B

Paper: https://arxiv.org/pdf/2607.25857

u/ai-lover — 13 days ago

NVIDIA AI Releases NOOA: An Object-Oriented Python Framework That Turns an AI Agent Into a Single Python Class

NVIDIA AI's NOOA: An Object-Oriented Python Framework That Turns an AI Agent Into a Single Python Class

Here's what's actually interesting:

  1. The whole agent is one classMethods are the actions the model can take. Fields are state. Docstrings are prompts. Type annotations are contracts the runtime enforces. A method whose body is ... becomes an LLM-driven loop; a method with a real body stays deterministic Python the model can call as a tool.

  2. Pass by reference is the load-bearing pieceArguments stay live in the execution environment. The model sees a bounded preview — concrete type, true length, head/tail sample — and writes code against the real object. → SWE-bench sessions peaked at 22–72k prompt tokens against 200–400k windows → No context compaction needed

  3. The benchmark numbers

→ 82.2% SWE-bench Verified with GPT-5.5, from a benchmark-agnostic 253-line agent

→ 86.8% CyberGym L1 with network access blocked, top open-source result reported

→ 85.1% mean RHAE on ARC-AGI-3 with GPT-5.6-sol, under $20 per game

→ ~1.1M tokens and ~28 model calls per task, against 2.2M and 66 for the compared harness

Full analysis: https://www.marktechpost.com/2026/08/07/nvidia-ai-releases-nooa-an-object-oriented-python-framework/

Paper: https://arxiv.org/pdf/2607.20709

Technical details: https://developer.nvidia.com/blog/six-agent-harness-capabilities-for-higher-model-performance/

Repo: https://github.com/NVIDIA-NeMo/labs-OO-Agents/tree/main

marktechpost.com
u/ai-lover — 14 days ago

Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot

Microsoft open sourced a unit-test agent that researches your repository before it writes a single test.

`code-testing-generator` ships in the MIT-licensed dotnet/skills repo. It is polyglot, and it does something most test generators skip: it proves the tests are worth keeping.

  1. It reads the repo first
    → Detects language, test framework, and existing conventions
    → Finds the real build and test commands
    → Confirms the repo's own test command actually discovers the new tests

  2. It scales the work to the request
    → Direct: single file, write and validate immediately
    → Single pass: one Research→Plan→Implement cycle
    → Iterative: repeat until the coverage target is met

  3. It checks its own tests before reporting done
    → Reasons about small mutations that should make the tests fail
    → Flags weak or missing assertions
    → Maps every requested scenario to a dedicated test

  4. The benchmark
    → 140/152 tasks vs 120/152 for stock GitHub Copilot, same model, same prompts
    → 79/89 vague prompts vs 59/89 — failures fell from 30 to 10
    → 61/63 detailed prompts for both, dead even
    → 15/15 on diff-targeted tasks vs 0/15

Full analysis: https://marktechpost.com/2026/08/06/microsoft-open-sources-code-testing-generator/

Repo: https://github.com/dotnet/skills/blob/main/plugins/dotnet-test/agents/code-testing-generator.agent.md

Technical details: https://devblogs.microsoft.com/dotnet/polyglot-unit-testing-agent/

marktechpost.com
u/ai-lover — 14 days ago
▲ 32 r/OpenSourceeAI+1 crossposts

Prime Intellect Releases Prime Agent: An Open-Source RLM Harness Where Sub-Agents Are Function Calls Inside Persistent IPython Kernel

Prime Intellect Releases Prime Agent: An Open-Source RLM Harness Where Sub-Agents Are Function Calls Inside Persistent IPython Kernel

Most coding harnesses hand the model a fixed set of tools. Prime Agent hands it one: a persistent IPython kernel. Everything else — file edits, shell, sub-agents, compaction — is a function call inside that kernel.

1. Sub-agents are function calls, not a special mode

→ rlm("sub-task") spawns a full child session with its own model, kernel, and history

→ It returns at admission, not with the answer, so the parent never blocks

→ Replies arrive later through agent_message

→ Messaging is scoped to parent, sibling, or child only

→ Idle sub-agents leave memory after 30 minutes, then reload from disk when addressed

2. The harness edits itself

→ Harness state is formalized as H = (ρ, G, K, M): prompt, sub-agents, skills, memory

→ /refine reads the trajectory and applies the smallest relevant edit

→ Each refinement records its trigger and its outcome

→ The base system prompt stays immutable; bad updates roll back by ID

3. The benchmark numbers

→ 95.5% RHAE Best@1 on ARC-AGI-3 with Opus 5, above the reported human expert baseline of 95.4%

→ Three runs: 95.0, 95.2, 95.5

→ 99.97% Best@3, all 183/183 levels complete

→ Long-context suite: with open-weights GLM-5.2, Prime Agent beats Pi-mono on 8 of 9 evals

Full analysis: https://www.marktechpost.com/2026/08/06/prime-intellect-releases-prime-agent/

GitHub Repo: https://github.com/PrimeIntellect-ai/prime-agent

Technical details: https://www.primeintellect.ai/blog/prime-agent

u/ai-lover — 15 days ago

Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model

Meta Superintelligence Labs released Muse Code (in beta mode), a terminal coding agent in beta, powered by its new Muse Spark 1.2 model.

Here are some key takeaways:

1. Async background agents that outlive the task

Muse Code runs a simple agent loop plus a set of specialized background agents. These stay active throughout the session instead of being spawned per task.

→ Meta says this avoids redundant information gathering

→ The agents carry out next steps and choose when to report back to the main agent

→ Stated effect: lower latency and less steering on multi-step tasks

2. An append-only event log as the single source of truth

Every model call, tool run, approval, and edit is appended to a local event log.

→ Meta calls the runtime replay-exact and restart-safe

→ After a crash, the agent resumes precisely where it stopped

→ This is what makes long-running tasks survive failures

3. Three bundled skills

→ /plan turns a task into an approval-gated plan

→ /grill stress-tests that plan until it holds up

→ /goal works toward completion of the specified objective

4. Muse Spark 1.2 was co-trained with the harness

Training included rejection-sampled harness trajectories and recipe optimizations for goals, compaction, and subagents. Meta also integrated the Muse Code toolset directly to maximize harness compatibility.

Long-horizon training covered whole-repository generation, large end-to-end projects, and auto-research.

5. The kernel optimization case study

Meta ran iterative GPU kernel optimization over 1,000+ tool calls, up to 24 hours per run.

→ Benchmarked on KDA and MLA kernels for NVIDIA Hopper GPUs

→ KDA baseline is the FLA Triton implementation, with third-party kernel libraries prohibited

→ MLA reference is PyTorch at batch size 1, 64 heads, sequence length 8192, latent dimension 512

→ For MLA, the model built a two-kernel Triton pipeline reusing the shared KV latent as both K and V

Full analysis: https://www.marktechpost.com/2026/08/05/meta-superintelligence-labs-releases-muse-code/

Technical details: https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2

Model: https://developer.meta.com/ai/models/muse-spark/

u/ai-lover — 16 days ago

NVIDIA Releases Alpamayo 2 Super: A 34B Open Vision-Language-Action Model for Robotaxis and Autonomous Driving Under OpenMDW-1.1

NVIDIA Releases Alpamayo 2 Super: A 34B Open Vision-Language-Action Model for Robotaxis and Autonomous Driving Under OpenMDW-1.1

Here are some key takeaways:

1. The architecture is split

→ 32B VLM backbone, built on Cosmos 3 Super Reasoner, post-trained with reinforcement learning

→ 2.3B diffusion-based action decoder

→ Roughly 3x the scale of the 10B Alpamayo 1 and Alpamayo 1.5

2. It ranks first on LingoQA

→ Lingo-Judge score of 79.2, first among nearly 40 models evaluated

→ +17.0 over Qwen2.5-VL 72B, +15.1 over Gemini 2.5 Pro, +23.2 over GPT-4o

→ Closed-loop AlpaSim score of 1.50 ± 0.13 across 910 NuRec scenarios

→ Open-loop minADE₆ of 0.911 m at 6.4s on 937 challenging samples

3. One pass produces five outputs

→ A trajectory: 64 waypoints from 0.1s to 6.4s, each with ego-frame XYZ and a 3x3 rotation matrix

→ A Chain-of-Causation trace explaining the decision

→ A meta-action such as yield, lane change or stop

→ Reasoning auto-labels for training and validation data

→ Visual question answering with 2D grounding

4. The training corpus

→ ~115,000 hours of multi-camera driving video with egomotion and trajectory annotations

→ ~3,700,000 Chain-of-Causation traces

→ Inputs are six cameras and four historical frames each in the validated public notebook profiles

Full analysis: https://www.marktechpost.com/2026/08/05/nvidia-alpamayo-2-super-open-vla-model-autonomous-driving/

Model weights: https://huggingface.co/nvidia/Alpamayo2-Super

u/ai-lover — 16 days ago
▲ 13 r/AIDeveloperNews+2 crossposts

CopilotKit Open Sources Channels SDK: An MIT Licensed Library That Runs Any AG-UI Agent Inside Slack And Microsoft Teams

CopilotKit Open Sources Channels SDK: An MIT Licensed Library That Runs Any AG-UI Agent Inside Slack And Microsoft Teams

No per-platform rewrite. No platform credentials in your agent process. No second agent to maintain.

Here's how it works:

  1. Describe once, render native One message description is lowered to a serializable intermediate representation, then rendered in each platform's own format. → Block Kit on Slack, Adaptive Cards on Teams

  2. Your agent doesn't move It connects over AG-UI, so the model, tools and business logic stay where they are. → LangGraph, CrewAI, Mastra, Pydantic AI, Google ADK

  3. The runtime owns the lifecycle There is no channel.start(). You await channels.ready(), so a broken config fails startup loudly instead of silently. → ready() · status() · stop()

  4. The concurrency trap Turns default to "parallel", and only the managed adapter serializes same-thread deliveries. On a direct adapter, one shared agent instance means two runs corrupt each other. → "parallel" (default) · "serial" · "drop"

  5. The numbers → 0.7.3, shipped August 4, MIT licensed → 5 adapters: /slack, /teams, /discord, /telegram, /whatsapp → Node.js 22+, ESM only, one long-running process → Slack and Teams GA; Discord and WhatsApp next

The key takeaway: one agent, five adapters, and platform credentials that never touch your process. Every channel needs a CopilotKit Intelligence key — free tier included, no standalone path.

Full analysis: https://www.marktechpost.com/2026/08/04/copilotkit-open-sources-channels-sdk/

GitHub Repo: https://github.com/CopilotKit/channels-sdk

Technical details: https://www.copilotkit.ai/blog/channels-sdk

marktechpost.com
u/ai-lover — 16 days ago

Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

Cursor Just Open-Sourced Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

No CPU-GPU synchronization. No separate communication library.

Here's what's interesting:

1. Communication direction is a per-operation choice

Most implementations push tokens to the GPUs that need them. Cursor benchmarked both directions and split the decision.

→ Pull dispatch signalling: 18 µs, against 103 µs for push

→ Up to 29% higher NVLink utilization under expert imbalance

2. One schedule table, four operations

Pull-based forward dispatch, push-based forward combine, pull reverse-combine, push reverse-dispatch. Build the schedule once, reuse it everywhere.

→ Under 3% of total MoE runtime, device-side, no CPU round trip

3. Overlap granularity has an interior optimum

Too fine and the tensor cores stall at barriers. Too coarse and they sit waiting for the first tokens to land. The heuristic targets two full SM waves per expert-grouped GEMM.

→ 2,368-token minibatch floor for Kimi 2.5 shapes

4. A ring buffer removes the CPU from the loop

The usual fixes are dropping tokens or asking the CPU to size the buffers. MoK cycles a fixed few-hundred-megabyte ring at minibatch granularity instead, and walks it in reverse to cut activation replay in the backward pass.

→ Zero tokens dropped, zero CPU-GPU synchronization

5. The numbers

Layer benchmarks, single NVL72 rack, EP degree 64, against the fastest public baseline:

→ 2.37× MXFP8 forward, 1.92× BF16 forward

→ 1.78× MXFP8 backward, 1.58× BF16 backward

End-to-end, 512 GPUs across several GB300 NVL72 racks:

→ 760.9 → 1,070.2 tokens/sec/GPU, a 1.41× gain

Full analysis: https://www.marktechpost.com/2026/08/04/cursor-open-sources-mixture-of-kittens-mok-a-deterministic-moe-training-megakernel-for-gb300-nvl72-racks/

GitHub Repo: https://github.com/cursor/mixture-of-kittens

Technical details: https://cursor.com/blog/mixture-of-kittens

u/ai-lover — 17 days ago
▲ 18 r/AIDeveloperNews+1 crossposts

Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

Cursor Just Open-Sourced Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

No CPU-GPU synchronization. No separate communication library.

Here's what's interesting:

1. Communication direction is a per-operation choice

Most implementations push tokens to the GPUs that need them. Cursor benchmarked both directions and split the decision.

→ Pull dispatch signalling: 18 µs, against 103 µs for push

→ Up to 29% higher NVLink utilization under expert imbalance

2. One schedule table, four operations

Pull-based forward dispatch, push-based forward combine, pull reverse-combine, push reverse-dispatch. Build the schedule once, reuse it everywhere.

→ Under 3% of total MoE runtime, device-side, no CPU round trip

3. Overlap granularity has an interior optimum

Too fine and the tensor cores stall at barriers. Too coarse and they sit waiting for the first tokens to land. The heuristic targets two full SM waves per expert-grouped GEMM.

→ 2,368-token minibatch floor for Kimi 2.5 shapes

4. A ring buffer removes the CPU from the loop

The usual fixes are dropping tokens or asking the CPU to size the buffers. MoK cycles a fixed few-hundred-megabyte ring at minibatch granularity instead, and walks it in reverse to cut activation replay in the backward pass.

→ Zero tokens dropped, zero CPU-GPU synchronization

5. The numbers

Layer benchmarks, single NVL72 rack, EP degree 64, against the fastest public baseline:

→ 2.37× MXFP8 forward, 1.92× BF16 forward

→ 1.78× MXFP8 backward, 1.58× BF16 backward

End-to-end, 512 GPUs across several GB300 NVL72 racks:

→ 760.9 → 1,070.2 tokens/sec/GPU, a 1.41× gain

Full analysis: https://www.marktechpost.com/2026/08/04/cursor-open-sources-mixture-of-kittens-mok-a-deterministic-moe-training-megakernel-for-gb300-nvl72-racks/

GitHub Repo: https://github.com/cursor/mixture-of-kittens

Technical details: https://cursor.com/blog/mixture-of-kittens

u/ai-lover — 17 days ago
▲ 18 r/AIDeveloperNews+2 crossposts

Reflex Open Sources XY: A Rust-Backed Super-Fast Python Charting Library That Keeps 100 Million Point Charts Interactive

Reflex AI Open Sources XY: A Rust-Backed Super-Fast Python Charting Library That Keeps 100 Million Point Charts Interactive

Here are some key points:

1. The benchmark

→ 0.071 s at 10,000 points

→ 0.081 s at 100 million points

→ Matplotlib reaches 13.385 s at 50M, then does not render 100M

→ Plotly reaches 9.794 s at 25M, then does not render 50M

2. Why it stays flat

Most Python charting stacks create one drawable object per row. XY draws what the screen can actually show. M4 decimation starts above 10,000 rows on lines. Automatic scatter density starts above 200,000 points.

3. Export size

→ A 10-million-point interactive scatter exports to 258 KiB of HTML

→ The Plotly equivalent is 259 MiB

Apache-2.0, Python 3.11+, pip install xy.

Full analysis: https://www.marktechpost.com/2026/08/04/reflex-open-sources-xy-a-rust-backed-super-fast-python-charting-library-that-keeps-100-million-point-charts-interactive/

GitHub Repo: https://github.com/reflex-dev/xy

Technical details: https://reflex.dev/blog/xy-python-charting-library/

u/ai-lover — 17 days ago
▲ 6 r/AIDeveloperNews+1 crossposts

How to Secure AI Agents, MCP Servers, and LLM Apps in Production

How to Secure AI Agents, MCP Servers, and LLM Apps in Production

Application security rests on one assumption: software does what its code says.

---AI agents broke it.

Mend.io's new practitioner guide — 𝘚𝘦𝘤𝘶𝘳𝘪𝘯𝘨 𝘈𝘐 𝘢𝘨𝘦𝘯𝘵𝘴, 𝘔𝘊𝘗 𝘴𝘦𝘳𝘷𝘦𝘳𝘴 & 𝘓𝘓𝘔 𝘢𝘱𝘱𝘴 — starts from that break. An agent's behavior emerges from the model, the system prompt, retrieved context, and the tools it's permitted to call. The failure modes never appear in a CVE feed: prompt injection through data, over-permissioned agents causing damage without a single exploit, poisoned tool descriptions on MCP servers, EOL models serving predictions after patching stops.

The guide's answer is three moves:

𝗦𝗲𝗲: Inventory the agentic attack surface across five layers — interaction, agent, integration, model, code. Hunt shadow agents via repo signatures and network egress. Run every agent through a 12-point misconfiguration checklist.

𝗙𝗶𝘅: Enrich → prioritize → triage. Rank by reachability and agentic amplification, not severity scores. Automate FP closures only with evidence trails. Risk acceptance is never automated.

𝗣𝗿𝗼𝘁𝗲𝗰𝘁: Guardrails on every input and output — embedded Python SDK or standalone Docker API server. Inbound: injection patterns, jailbreaks. Outbound: credentials, PII, policy violations. The core design principle: an agent that can't call a dangerous tool doesn't need a prompt begging it not to.

Includes a 15-question maturity self-assessment aligned to NIST AI RMF, OWASP AIMA, ISO/IEC 42001, and the EU AI Act.

Full analysis: https://www.marktechpost.com/2026/08/03/how-to-secure-ai-agents-mcp-servers-and-llm-apps-in-production/

Download the full guide, free: https://pxllnk.co/lxn88m

pxllnk.co
u/ai-lover — 18 days ago

Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date

Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date

1. What shipped

→ 2.4T parameters, mixture-of-experts

→ 1M context, 991K max input, 131K max output

→ Text, image and video input

→ $2.00 input, $6.00 output, $0.25 cached input per 1M tokens

→ Open weights next week

2. Where it leads Fable5

→ Terminal Bench 2.1: 86.6 vs 84.6

→ PaperBench: 93.0 vs 88.8

→ IFBench: 82.8 vs 63.5

→ Parametric CAD Bench: 91.5 vs 87.5

→ OmniDocBench 1.5: 92.1 vs 89.5

3. Where it trails Fable5

→ SWE-bench Pro: 67.7 vs 80.0

→ FrontierSWE: 73.5 vs 88.8

→ HLE: 43.6 vs 53.3

→ Toolathlon Verified: 72.5 vs 77.9

4. The category split

→ Multimodal Reasoning: above Fable5 on 11 of 11 rows

→ Document & Office: 7 of 7

→ Perception & Grounding: 9 of 10

→ Coding Agent: 3 of 11

→ General Agent: 1 of 8

→ Visual Agent & Coding: 3 of 11

Full analysis: https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/

Technical details: https://qwen.ai/blog?id=qwen3.8

API: https://www.qwencloud.com/models/qwen3.8-max#context

u/ai-lover — 18 days ago
▲ 56 r/OpenSourceeAI+1 crossposts

NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework

NVIDIA released Molt, a PyTorch-native training framework for agentic reinforcement learning.

Here is what stands out technically:

1. The footprint is the feature

→ ~8.6K lines of RL code, counted by tracing the import graph from the RL entry point

→ Same method: ~62K for verl, ~25K for slime, ~7.2K for OpenRLHF

→ One training backend (NeMo AutoModel), one serving engine (vLLM), neither forked

2. Three components, one asynchronous loop

→ Ray for placement and the async queue, vLLM for rollout, FSDP2 + AutoModel for a single trainable actor → A streaming pool keeps prompt groups in flight so engines never drain while the actor trains

→ Partial rollout pauses engines, broadcasts shards over NCCL, and resumes retained requests instead of discarding them

3. The agent is an ordinary Python program

→ One module exporting an AgentRunner; reward is any Python you write

→ Env gives you a Gymnasium-style step(); ChatAgent lets a stock OpenAI or Anthropic SDK train as-is

→ A loopback server captures token ids and log-probabilities, so retokenization drift never enters the trajectory

Full analysis: https://www.marktechpost.com/2026/08/01/nvidia-ai-releases-molt-a-pytorch-native-agentic-reinforcement-learning-framework/

Paper: https://arxiv.org/pdf/2607.21653

Repo: https://github.com/NVIDIA-NeMo/labs-molt

u/ai-lover — 19 days ago