Morning Singularity Digest - 2026-07-23

Estimated total read • ~35 min

Skim fast, dive deep only where it matters.

2-minute skim 10-minute read Deep dive optional
Contents

Front Page

~10 min

nexu-io/open-design: 🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.

Signal 10.0 Novelty 7.3 Impact 7.7 Confidence 7.0 Actionability 6.5

Summary: 🎨 The open-source Claude Design alternative.

  • What happened: 🎨 The open-source Claude Design alternative.
  • Why it matters: 0.13.0 keeps the session alive: resume Codex / OpenCode / Pi / Open Design Cloud runs across turns, pick the right model faster, and hand off screenshot-backed PPTX /.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

🎨 The open-source Claude Design alternative.

What's new

🖥️ Local-first native desktop app for macOS and Windows.

Key details

  • 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export.
  • 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • 🔥 Open Design 0.13.0 — Stay in Flow is here.
  • Long design sessions used to break on every interruption — a run lost its place, a model picker made you guess, an export needed one more detour.

Results & evidence

  • 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • 🔥 Open Design 0.13.0 — Stay in Flow is here.
  • 0.13.0 keeps the session alive: resume Codex / OpenCode / Pi / Open Design Cloud runs across turns, pick the right model faster, and hand off screenshot-backed PPTX / PDF without leaving the app.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Signal 10.0 Novelty 6.2 Impact 8.3 Confidence 7.0 Actionability 6.5

Summary: The agent harness performance optimization system.

  • What happened: The agent harness performance optimization system.
  • Why it matters: The agent harness performance optimization system.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

The agent harness performance optimization system.

What's new

Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Key details

  • Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
  • Language: English | Português (Brasil) | 简体中文 | 繁體中文 | 日本語 | 한국어 | Türkçe | Русский | Tiếng Việt | ไทย | Deutsch | Español Warning Official sources only.
  • Install ECC only from verified channels: the GitHub repository github.com/affaan-m/ECC, the npm packages ecc-universal and ecc-agentshield, the GitHub App, the plugin slug ecc@ecc, and the project website ecc.tools.
  • Third-party re-uploads and unofficial mirrors are not maintained or reviewed by the project and may contain malware.

Results & evidence

  • 211.9K+ stars | 32.5K+ forks | 230+ contributors | 12+ language ecosystems | Cross-harness agent workflows Language / 语言 / 語言 / Dil / Язык / Ngôn ngữ / Idioma English | Português (Brasil) | 简体中文 | 繁體中文 | 日本語 | 한국어 | Türkçe | Русский | Tiếng Việt | ไทย | Deu...
  • Production-ready agents, skills, hooks, rules, MCP configurations, and legacy command shims evolved over 10+ months of intensive daily use building real products.
  • ECC v2.0.0 adds the public Hermes operator story on top of that reusable layer: start with the Hermes setup guide, then review the 2.0.0 release notes and cross-harness architecture.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level.

  • What happened: We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
  • Why it matters: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...

What's new

arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.

Key details

  • However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
  • This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
  • We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
  • SHOVIR extends two spatially annotated chest X-ray datasets, MIMIC-CXR and PadChest-GR, with per-box CheXpert labels, and defines image-level and disease-level occlusion experiments that contrast baseline performance on clean images against localized, regio...

Results & evidence

  • arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.
  • Computer Science > Computer Vision and Pattern Recognition [Submitted on 29 Jun 2026 (v1), last revised 22 Jul 2026 (this version, v2)] Title:SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation View PDF HTML (experimen...
  • Submission history From: Filippo Ruffini [view email][v1] Mon, 29 Jun 2026 12:17:35 UTC (1,063 KB) [v2] Wed, 22 Jul 2026 13:09:26 UTC (1,063 KB) References & Citations Loading...

Limitations / unknowns

  • However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
  • This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
  • Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 8.7 Actionability 6.5

Summary: arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex.

  • What happened: This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation.
  • Why it matters: arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Submission history From: Bingru Li [view email][v1] Thu, 5 Feb 2026 09:55:19 UTC (1,414 KB) [v2] Tue, 21 Jul 2026 17:59:46 UTC (1,463 KB) Current browse context: cs.CL References & Citations Loading...

What's new

arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.

Key details

  • While Large Language Models (LLMs) show promise, a significant gap remains between the theoretical capability of LLMs and their practical utility for researchers.
  • This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation.
  • The platform comprises an Annotator and an optional Reviewer to simulate a peer-review process.
  • This platform supports comparative experiments across three main paradigms: Prompt Engineering (Zero-shot/Few-shot/Chain-of-thought), Retrieval-Augmented Generation, and Fine-tuning.

Results & evidence

  • arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.
  • Computer Science > Computation and Language [Submitted on 5 Feb 2026 (v1), last revised 21 Jul 2026 (this version, v2)] Title:LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation View PDF HTML (experimental)A...
  • Submission history From: Bingru Li [view email][v1] Thu, 5 Feb 2026 09:55:19 UTC (1,414 KB) [v2] Tue, 21 Jul 2026 17:59:46 UTC (1,463 KB) Current browse context: cs.CL References & Citations Loading...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Mwe-MCP – self-hosted memory for AI agents that knows who may know what

Signal 8.4 Novelty 5.1 Impact 2.6 Confidence 7.5 Actionability 3.5

Summary: Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that behaves like a.

  • What happened: Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that.
  • Why it matters: Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that behaves like a member of the family, not as a personal agent, and this made me clash very ear...

What's new

First on OpenClaw, as a plugin, then the idea matured and since the beginning of this year the memory plugin has evolved into an agent agnostic MCP server.

Key details

  • First on OpenClaw, as a plugin, then the idea matured and since the beginning of this year the memory plugin has evolved into an agent agnostic MCP server.
  • It has been running my household since spring, 4 people and 3 agents on the same memory.

    The core idea that sets this project apart is the addition of inline ACLs on the prose of the wiki pages; the inspiration came to me, in all honesty, when social medi...

  • In practice: my partner can tell the assistant something and our kid can't get it out of it, "we're out of milk" belongs to the whole household and not to me; the same page is served differently to each reader, redacted by the engine bef...
  • Memory is organized into wikis, every user, group and agent has their own wiki and other wikis can emerge autonomously.

    To reach this goal the server uses various components that use configurable LLMs, from the simple dedup, which you can easily run local...

Results & evidence

  • It has been running my household since spring, 4 people and 3 agents on the same memory.

    The core idea that sets this project apart is the addition of inline ACLs on the prose of the wiki pages; the inspiration came to me, in all honesty, when social medi...

Limitations / unknowns

  • The prose isn't there for beauty or for humans to read, the prose links the facts, so that no fact is disconnected from the rest but everything is part of one huge narrative, which is however made extremely efficient by the metadata.

    These are the pe...

  • Now I've realized that this memory system, initially born for a family, can be great for a team too and, with some future adjustments, even for small and medium businesses.

    Honest limits: the internal LLM needs a reasonably capable model (small local...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

What Changed Overnight

~1 min
  • New: SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation
  • New: New Framework Desktop Option with AMD Ryzen AI Max+ Pro 495 and 192GB Memory
  • New: Solar Open 2 Technical Report
  • New: The Two-Process Theory of Machine Self-Report
  • New: DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
  • New: OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
  • Removed: PathReportEval: A Systematic Benchmark for Pathology Report Generation (fell below rank threshold)
  • Removed: Uncertainty Quantification for AI-Driven Crash Simulation Surrogates: A Comparative Study of Monte Carlo Dropout and Deep Ensemble on Open-Source Bumper Beam Benchmark (fell below rank threshold)
  • Removed: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack (fell below rank threshold)
  • Removed: AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents (fell below rank threshold)
  • What to do now:
  • Validate with one small internal benchmark and compare against your current baseline this week.
  • Track for corroboration and benchmark data before adopting.

Deep Dives

~7 min

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level.

  • What happened: We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
  • Why it matters: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...

What's new

arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.

Key details

  • However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
  • This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
  • We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
  • SHOVIR extends two spatially annotated chest X-ray datasets, MIMIC-CXR and PadChest-GR, with per-box CheXpert labels, and defines image-level and disease-level occlusion experiments that contrast baseline performance on clean images against localized, regio...

Results & evidence

  • arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.
  • Computer Science > Computer Vision and Pattern Recognition [Submitted on 29 Jun 2026 (v1), last revised 22 Jul 2026 (this version, v2)] Title:SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation View PDF HTML (experimen...
  • Submission history From: Filippo Ruffini [view email][v1] Mon, 29 Jun 2026 12:17:35 UTC (1,063 KB) [v2] Wed, 22 Jul 2026 13:09:26 UTC (1,063 KB) References & Citations Loading...

Limitations / unknowns

  • However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
  • This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
  • Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Mwe-MCP – self-hosted memory for AI agents that knows who may know what

Signal 8.4 Novelty 5.1 Impact 2.6 Confidence 7.5 Actionability 3.5

Summary: Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that behaves like a.

  • What happened: Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that.
  • Why it matters: Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Hi HN, for about a year now I've been experimenting with AI agents and building my own home ecosystem; from the start I set out with the idea of an agent that behaves like a member of the family, not as a personal agent, and this made me clash very ear...

What's new

First on OpenClaw, as a plugin, then the idea matured and since the beginning of this year the memory plugin has evolved into an agent agnostic MCP server.

Key details

  • First on OpenClaw, as a plugin, then the idea matured and since the beginning of this year the memory plugin has evolved into an agent agnostic MCP server.
  • It has been running my household since spring, 4 people and 3 agents on the same memory.

    The core idea that sets this project apart is the addition of inline ACLs on the prose of the wiki pages; the inspiration came to me, in all honesty, when social medi...

  • In practice: my partner can tell the assistant something and our kid can't get it out of it, "we're out of milk" belongs to the whole household and not to me; the same page is served differently to each reader, redacted by the engine bef...
  • Memory is organized into wikis, every user, group and agent has their own wiki and other wikis can emerge autonomously.

    To reach this goal the server uses various components that use configurable LLMs, from the simple dedup, which you can easily run local...

Results & evidence

  • It has been running my household since spring, 4 people and 3 agents on the same memory.

    The core idea that sets this project apart is the addition of inline ACLs on the prose of the wiki pages; the inspiration came to me, in all honesty, when social medi...

Limitations / unknowns

  • The prose isn't there for beauty or for humans to read, the prose links the facts, so that no fact is disconnected from the rest but everything is part of one huge narrative, which is however made extremely efficient by the metadata.

    These are the pe...

  • Now I've realized that this memory system, initially born for a family, can be great for a team too and, with some future adjustments, even for small and medium businesses.

    Honest limits: the internal LLM needs a reasonably capable model (small local...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically

Signal 10.0 Novelty 5.1 Impact 7.8 Confidence 7.0 Actionability 6.5

Summary: AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other.

  • What happened: AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping.
  • Why it matters: It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Instead, you are programming the program.md Markdown files that provide context to the AI agents and set up your autonomous research org.

What's new

AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ri...

Key details

  • Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies.
  • The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension.
  • This repo is the story of how it all began.
  • The idea: give an AI agent a small but real LLM training setup and let it experiment autonomously overnight.

Results & evidence

  • The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension.
  • It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Reality Check

~1 min
  • nexu-io/open-design: 🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • Primary source: yes
  • Demo available: yes
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • Show HN: Mwe-MCP – self-hosted memory for AI agents that knows who may know what
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • Show HN: Mwe-MCP – self-hosted memory for AI agents that knows who may know what
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.

Lab Notes

~1 min
  • Tool/Repo of the day: nexu-io/open-design: 🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK. (https://github.com/nexu-io/open-design)
  • Prompt/Workflow of the day: summarize claim -> evidence -> risk in three passes before acting.
  • Tiny snippet: `uv run python -m msd.run --scheduled`

Research Radar

~6 min

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level.

  • What happened: We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
  • Why it matters: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...

What's new

arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.

Key details

  • However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
  • This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
  • We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
  • SHOVIR extends two spatially annotated chest X-ray datasets, MIMIC-CXR and PadChest-GR, with per-box CheXpert labels, and defines image-level and disease-level occlusion experiments that contrast baseline performance on clean images against localized, regio...

Results & evidence

  • arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.
  • Computer Science > Computer Vision and Pattern Recognition [Submitted on 29 Jun 2026 (v1), last revised 22 Jul 2026 (this version, v2)] Title:SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation View PDF HTML (experimen...
  • Submission history From: Filippo Ruffini [view email][v1] Mon, 29 Jun 2026 12:17:35 UTC (1,063 KB) [v2] Wed, 22 Jul 2026 13:09:26 UTC (1,063 KB) References & Citations Loading...

Limitations / unknowns

  • However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
  • This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
  • Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 8.7 Actionability 6.5

Summary: arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex.

  • What happened: This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation.
  • Why it matters: arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Submission history From: Bingru Li [view email][v1] Thu, 5 Feb 2026 09:55:19 UTC (1,414 KB) [v2] Tue, 21 Jul 2026 17:59:46 UTC (1,463 KB) Current browse context: cs.CL References & Citations Loading...

What's new

arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.

Key details

  • While Large Language Models (LLMs) show promise, a significant gap remains between the theoretical capability of LLMs and their practical utility for researchers.
  • This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation.
  • The platform comprises an Annotator and an optional Reviewer to simulate a peer-review process.
  • This platform supports comparative experiments across three main paradigms: Prompt Engineering (Zero-shot/Few-shot/Chain-of-thought), Retrieval-Augmented Generation, and Fine-tuning.

Results & evidence

  • arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.
  • Computer Science > Computation and Language [Submitted on 5 Feb 2026 (v1), last revised 21 Jul 2026 (this version, v2)] Title:LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation View PDF HTML (experimental)A...
  • Submission history From: Bingru Li [view email][v1] Thu, 5 Feb 2026 09:55:19 UTC (1,414 KB) [v2] Tue, 21 Jul 2026 17:59:46 UTC (1,463 KB) Current browse context: cs.CL References & Citations Loading...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.7 Actionability 6.5

Summary: arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and.

  • What happened: We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation.
  • Why it matters: arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning.

What's new

arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user i...

Key details

  • It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video.
  • Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response.
  • We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p.
  • Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts.

Results & evidence

  • arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user i...
  • We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p.
  • We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Forecast & Watchlist

~1 min
  • Watch: agent
  • Watch: llm
  • Watch: cs.ai
  • Watch: cs.lg
  • Watch: rss
  • Watch: cs.cl
  • Watch: python
  • Watch: benchmark

Save for Later

~8 min

mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.

Signal 10.0 Novelty 5.1 Impact 8.2 Confidence 7.0 Actionability 6.5

Summary: Straight from my .agents directory.

  • What happened: Straight from my .agents directory.
  • Why it matters: Straight from my .agents directory.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Straight from my .agents directory.

What's new

Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.

Key details

  • My agent skills that I use every day to do real engineering - not vibe coding.
  • Developing real applications is hard.
  • Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.
  • But while doing so, they take away your control and make bugs in the process hard to resolve.

Results & evidence

  • If you want to keep up with changes to these skills, and any new ones I create, you can join ~60,000 other devs on my newsletter: - Run the skills.sh installer: npx skills@latest add mattpocock/skills- Pick the skills you want, and which coding agents you w...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

ultraworkers/claw-code: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Signal 10.0 Novelty 5.1 Impact 8.2 Confidence 7.0 Actionability 6.5

Summary: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

  • What happened: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
  • Why it matters: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

For file submission/navigation questions, see Navigation and file context.

What's new

Windows users can jump to the PowerShell-first Windows install and release quickstart.

Key details

  • github.com/code-yeongyu/lazycodex github.com/Yeachan-Heo/gajae-code Join the Discords: ultraworkers discord · gajae-code discord Important Claw Code is not the serious production project here.
  • This repository is closer to a museum exhibit than a product pitch, a crustacean-run artifact kept alive by clawed gajaes, swept and labeled by agents, and automatically maintained according to the harnesses above.
  • As already described in the project philosophy, this is not meant to be hand-operated like a normal product repo.
  • It is an agent-managed exhibit: the harnesses plan, execute, verify, label, and preserve the artifact while the crabs keep the tank running.

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.7 Actionability 6.5

Summary: arXiv:2607.18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge.

  • What happened: arXiv:2607.18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world.
  • Why it matters: arXiv:2607.18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

arXiv:2607.18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge.

What's new

arXiv:2607.18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge.

Key details

  • As embodied agents become increasingly capable, there is a growing demand for compact models that can serve as an on-device brain, preserving the broad general intelligence of LLMs while enabling effective high-level interaction with embodied environments.
  • Existing approaches, however, often prioritize either general-purpose intelligence or specialized embodied capabilities, making it challenging to satisfy both requirements within a single model.
  • We present \textbf{Athena-Brain-8B}, an 8B LLM designed to serve as an on-device brain for embodied intelligence for embodied intelligence.
  • Through a multi-stage post-training pipeline consisting of General Supervised Fine-Tuning, General Reinforcement Learning, Embodied Expert training, and Model Merge, Athena-Brain-8B maintains strong general capabilities while acquiring strong high-level emb...

Results & evidence

  • arXiv:2607.18985v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge.
  • Computer Science > Artificial Intelligence [Submitted on 21 Jul 2026] Title:Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio View PDF HTML (experimental)Abstract:Large language models (LLMs) have demon...

Limitations / unknowns

  • Existing approaches, however, often prioritize either general-purpose intelligence or specialized embodied capabilities, making it challenging to satisfy both requirements within a single model.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Nova – open-source AI orchestrator that works with you

Signal 8.4 Novelty 5.1 Impact 2.6 Confidence 7.5 Actionability 3.5

Summary: Show HN: Nova – open-source AI orchestrator that works with you

  • What happened: Show HN: Nova – open-source AI orchestrator that works with you
  • Why it matters: Could materially affect near-term AI workflows.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Show HN: Nova – open-source AI orchestrator that works with you

What's new

Show HN: Nova – open-source AI orchestrator that works with you

Key details

  • Show HN: Nova – open-source AI orchestrator that works with you

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

New Framework Desktop Option with AMD Ryzen AI Max+ Pro 495 and 192GB Memory

Signal 8.6 Novelty 5.1 Impact 5.2 Confidence 6.2 Actionability 3.5

Summary: Framework Desktop is a big computer made mini.

  • What happened: Framework Desktop is a big computer made mini.
  • Why it matters: Framework Desktop is a big computer made mini.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Framework Desktop is a big computer made mini.

What's new

Framework Desktop is a big computer made mini.

Key details

  • Massive gaming capability, heavy-duty AI compute, and standard PC parts, all in 4.5L.
  • Seize the means of computation The most powerful Framework Desktop yet is coming soon with an AMD Ryzen™ AI Max+ PRO 495 processor and 192GB of LPDDR5X memory.
  • 192GB Unified memory 273GB/s Memory bandwidth 131 TOPS AI compute 16C/32T Zen 5 CPU Linux Optionally pre-loaded 40-CU AMD Radeon™ 8065S GPU

Results & evidence

  • Massive gaming capability, heavy-duty AI compute, and standard PC parts, all in 4.5L.
  • Seize the means of computation The most powerful Framework Desktop yet is coming soon with an AMD Ryzen™ AI Max+ PRO 495 processor and 192GB of LPDDR5X memory.
  • 192GB Unified memory 273GB/s Memory bandwidth 131 TOPS AI compute 16C/32T Zen 5 CPU Linux Optionally pre-loaded 40-CU AMD Radeon™ 8065S GPU

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Understanding the AI Economy

Signal 8.6 Novelty 4.0 Impact 5.2 Confidence 6.2 Actionability 3.5

Summary: Understanding the AI economy There is broad agreement that AI’s potential to transform the global economy and the way we work is significant.

  • What happened: Understanding the AI economy There is broad agreement that AI’s potential to transform the global economy and the way we work is significant.
  • Why it matters: To get there, we as a society must work together to positively shape how AI impacts our lives, jobs, and economy.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Understanding the AI economy There is broad agreement that AI’s potential to transform the global economy and the way we work is significant.

What's new

To help, Google is launching the first iteration of the AI & Economy ATLAS (Activity, Task, Landscape, and Adoption Study), an ongoing, large-scale, de-identified study of how people are using Google’s AI products and tools.

Key details

  • However, the outcomes – what this means for work, for people’s lives, and the economy writ large – are not automatic nor guaranteed.
  • To get there, we as a society must work together to positively shape how AI impacts our lives, jobs, and economy.
  • In order for this shared work to be effective, it is critical to have a rich understanding of how AI is being adopted and used in the economy.
  • Society needs empirical insights and evidence-based research to inform decisions, initiatives, and actions.

Results & evidence

  • ATLAS’s first dataset (v1.0) is built from 15 million aggregated and de-identified human-AI interactions across the Gemini App, AI Mode, and the Gemini API, which together are used by more than 1 billion people monthly.
  • ATLAS v1.0 insights span more than 150 countries, 140 languages, 800 occupations, and 4,000 tasks; ATLAS is the most comprehensive look to date at how real people are using AI at scale.
  • The ATLAS v1.0 report provides an early view of a quickly moving landscape: AI’s capabilities are advancing, its use is evolving, and tools for observing its impact on the economy are still a work-in-progress.

Limitations / unknowns

  • However, the outcomes – what this means for work, for people’s lives, and the economy writ large – are not automatic nor guaranteed.
  • However within jobs, people are using AI selectively: in a typical job AI is used for only ~21% of tasks.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.