Morning Singularity Digest - 2026-08-03

Estimated total read • ~33 min

Skim fast, dive deep only where it matters.

2-minute skim 10-minute read Deep dive optional
Contents

Front Page

~9 min

nexu-io/open-design: ๐ŸŽจ The open-source Claude Design alternative. ๐Ÿ–ฅ๏ธ Local-first desktop app. ๐Ÿ–ผ๏ธ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video โ€” real files, HTML/PDF/PPTX/MP4 export. ๐Ÿค– Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.

Signal 10.0 Novelty 7.3 Impact 7.8 Confidence 7.0 Actionability 6.5

Summary: ๐ŸŽจ The open-source Claude Design alternative.

  • What happened: ๐ŸŽจ The open-source Claude Design alternative.
  • Why it matters: ๐ŸŽจ The open-source Claude Design alternative.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

๐ŸŽจ The open-source Claude Design alternative.

What's new

๐Ÿ–ฅ๏ธ Local-first native desktop app for macOS and Windows.

Key details

  • ๐Ÿ–ผ๏ธ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video โ€” real files, HTML/PDF/PPTX/MP4 export.
  • ๐Ÿค– Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • โšก Open Design Cloud โ€” the official model service.
  • One recharge to use GPT, Claude, Gemini, and DeepSeek inside Open Design: 20+ flagship models, zero config, billed by real token usage.

Results & evidence

  • ๐Ÿค– Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • One recharge to use GPT, Claude, Gemini, and DeepSeek inside Open Design: 20+ flagship models, zero config, billed by real token usage.
  • ๐Ÿค– Runs on Claude Code ยท OpenClaw ยท Codex ยท Cursor ยท OpenCode ยท Qwen ยท Copilot ยท Amp ยท Hermes ยท Kimi ยท Antigravity and 25 distinct local CLI executables, or any OpenAI-compatible endpoint via BYOK.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Signal 10.0 Novelty 6.2 Impact 8.3 Confidence 7.0 Actionability 6.5

Summary: The agent harness performance optimization system.

  • What happened: The agent harness performance optimization system.
  • Why it matters: plan -> test -> implement -> review -> verify -> remember -> improve Instead of rebuilding that process in every prompt, you install it once and make it part of how your.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

The agent harness performance optimization system.

What's new

Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Key details

  • Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
  • Language: English | Portuguรชs (Brasil) | ็ฎ€ไฝ“ไธญๆ–‡ | ็น้ซ”ไธญๆ–‡ | ๆ—ฅๆœฌ่ชž | ํ•œ๊ตญ์–ด | Tรผrkรงe | ะ ัƒััะบะธะน | Tiแบฟng Viแป‡t | เน„เธ—เธข | Deutsch | Espaรฑol Warning Official sources only.
  • Install ECC only from verified channels: the GitHub repository github.com/affaan-m/ECC, the npm packages ecc-universal and ecc-agentshield, the GitHub App, the plugin slug ecc@ecc, and the project website ecc.tools.
  • Third-party re-uploads and unofficial mirrors are not maintained or reviewed by the project and may contain malware.

Results & evidence

  • | ECC Pro + GitHub App Install free ยท Private repos from $19/seat/mo | Sponsor ECC Fund the open-source project | Community Discord ยท Q&A ยท Show and Tell | OSS stays free.
  • That's why a single maintainer ships weekly across 7 harnesses.
  • Access to 67 agents, 281 skills, and 94 legacy command shims, plus hooks, rules, memory, continuous learning, and AgentShield security scanning.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.

  • What happened: We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single.
  • Why it matters: Under exclusion, accuracy over all cases is pinned down only to a worst-case interval as wide as the uncovered fraction.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Submission history From: Delip Rao [view email] [v1] Mon, 25 May 2026 07:31:44 UTC (269 KB) [v2] Fri, 31 Jul 2026 04:03:56 UTC (615 KB) Ancillary-file links: Ancillary files (details): Current browse context: cs.CL Change to browse by: References & Citation...

What's new

arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.

Key details

  • Yet the same verdicts can support wildly varying agreement numbers, depending on seemingly minor choices: the judgment scale, the retained cases, the handling of abstentions and invalid outputs, and the pooling of verdicts across items and rubric criteria.
  • The statistics that settle these choices are established, but in psychometrics, econometrics, and corpus annotation rather than in the evaluation practice that needs them.
  • We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single source-attributed analysis, and apply it to three published LLM-judge evaluations.
  • For non-degenerate binary verdicts, Pearson's $r$, Spearman's $\rho$, Kendall's $\tau_b$, the phi coefficient, and the Matthews correlation coefficient are exactly the same statistic, so reporting several repeats one number under different names.

Results & evidence

  • arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.
  • Cohen's $\kappa$ differs from them only through a marginal-mismatch factor in $(0,1]$ and shares their asymptotic variance when judge and human assign the positive verdict equally often.
  • On a rubric benchmark carrying per-criterion human labels, protocol choice alone moves reported accuracy from $0.551$ to $0.899$ and carries $\kappa$ across zero, without altering a single verdict.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema.

  • What happened: arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined.
  • Why it matters: We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios.

What's new

arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with sour...

Key details

  • We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together.
  • The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios.
  • The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms.
  • We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1.

Results & evidence

  • arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with sour...
  • The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios.
  • Computer Science > Artificial Intelligence [Submitted on 31 Jul 2026] Title:ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction View PDF HTML (experimental) Abstract:Enterprise workflows increasingly rely on agents for \emph{schema-gu...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Nightcrawler โ€“ A local AI pentesting agent running on a smartphone

Signal 8.5 Novelty 5.1 Impact 4.7 Confidence 7.5 Actionability 3.5

Summary: An autonomous penetration testing agent that runs entirely on a smartphone.

  • What happened: An autonomous penetration testing agent that runs entirely on a smartphone.
  • Why it matters: An autonomous penetration testing agent that runs entirely on a smartphone.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

An autonomous penetration testing agent that runs entirely on a smartphone.

What's new

- Fully autonomous โ€” no human in the loop during operation - 100% local inference โ€” AI runs on the phone's GPU, no cloud needed - Scope-enforced โ€” two-layer defense prevents out-of-scope actions - Stealth-first โ€” slow scan rates, host rotation, cover traffi...

Key details

  • Drop the phone on a network, walk away, and it discovers hosts, maps services, finds vulnerabilities, and generates a pentest report โ€” all without cloud connectivity.
  • โ–‘โ–ˆโ–„โ–‘โ–ˆ โ–ˆ โ–ˆโ–€โ–€ โ–ˆโ–‘โ–ˆ โ–€โ–ˆโ–€ โ–ˆโ–€โ–€ โ–ˆโ–€โ–ˆ โ–„โ–€โ–ˆ โ–ˆโ–‘โ–ˆโ–‘โ–ˆ โ–ˆโ–‘โ–‘ โ–ˆโ–€โ–€ โ–ˆโ–€โ–ˆ โ–‘โ–ˆโ–‘โ–€โ–ˆ โ–ˆ โ–ˆโ–„โ–ˆ โ–ˆโ–€โ–ˆ โ–‘โ–ˆโ–‘ โ–ˆโ–„โ–„ โ–ˆโ–€โ–„ โ–ˆโ–€โ–ˆ โ–€โ–„โ–€โ–„โ–€ โ–ˆโ–„โ–„ โ–ˆโ–ˆโ–„ โ–ˆโ–€โ–„ v0.1.0 AUTONOMOUS MOBILE PENTEST AGENT OnePlus 8 ยท NetHunter ยท LFM2.5-1.2B ยท OpenCL GPU Penetration testing (pentesting) is the practice of testing a compute...
  • Professional pentesters are hired to find vulnerabilities before real attackers do.
  • Nightcrawler automates this process on a phone.

Results & evidence

  • โ–‘โ–ˆโ–„โ–‘โ–ˆ โ–ˆ โ–ˆโ–€โ–€ โ–ˆโ–‘โ–ˆ โ–€โ–ˆโ–€ โ–ˆโ–€โ–€ โ–ˆโ–€โ–ˆ โ–„โ–€โ–ˆ โ–ˆโ–‘โ–ˆโ–‘โ–ˆ โ–ˆโ–‘โ–‘ โ–ˆโ–€โ–€ โ–ˆโ–€โ–ˆ โ–‘โ–ˆโ–‘โ–€โ–ˆ โ–ˆ โ–ˆโ–„โ–ˆ โ–ˆโ–€โ–ˆ โ–‘โ–ˆโ–‘ โ–ˆโ–„โ–„ โ–ˆโ–€โ–„ โ–ˆโ–€โ–ˆ โ–€โ–„โ–€โ–„โ–€ โ–ˆโ–„โ–„ โ–ˆโ–ˆโ–„ โ–ˆโ–€โ–„ v0.1.0 AUTONOMOUS MOBILE PENTEST AGENT OnePlus 8 ยท NetHunter ยท LFM2.5-1.2B ยท OpenCL GPU Penetration testing (pentesting) is the practice of testing a compute...
  • It uses a small AI model (LFM2.5-1.2B-Instruct-Heretic, 1.2 billion parameters) running locally on the phone's GPU to decide what to do next โ€” which host to probe, which tool to use, what to look for.
  • โ”‚ commands โ”‚ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ”‚ โ”‚ โ”‚ โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚ โ”‚ โ”‚ Web Dashboard โ”‚ โ”‚ SQLite DB โ”‚ โ”‚ โ”‚ โ”‚ (:8888) โ”‚ โ”‚ Hosts, vulns, โ”‚ โ”‚ โ”‚ โ”‚ Monitor & steer โ”‚ โ”‚ creds, commands โ”‚ โ”‚ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

What Changed Overnight

~1 min
  • New: nexu-io/open-design: ๐ŸŽจ The open-source Claude Design alternative. ๐Ÿ–ฅ๏ธ Local-first desktop app. ๐Ÿ–ผ๏ธ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video โ€” real files, HTML/PDF/PPTX/MP4 export. ๐Ÿค– Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • New: mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.
  • New: karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically
  • New: multica-ai/andrej-karpathy-skills: A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
  • New: Show HN: Nightcrawler โ€“ A local AI pentesting agent running on a smartphone
  • New: Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
  • Removed: paperclipai/paperclip: The open-source app everyone uses to manage agents at work (fell below rank threshold)
  • Removed: Panniantong/Agent-Reach: Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu โ€” one CLI, zero API fees. (fell below rank threshold)
  • Removed: colbymchenry/codegraph: Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent โ€” fewer tokens, fewer tool calls, 100% local (fell below rank threshold)
  • Removed: rtk-ai/rtk: CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies (fell below rank threshold)
  • What to do now:
  • Validate with one small internal benchmark and compare against your current baseline this week.
  • Track for corroboration and benchmark data before adopting.

Deep Dives

~7 min

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.

  • What happened: We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single.
  • Why it matters: Under exclusion, accuracy over all cases is pinned down only to a worst-case interval as wide as the uncovered fraction.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Submission history From: Delip Rao [view email] [v1] Mon, 25 May 2026 07:31:44 UTC (269 KB) [v2] Fri, 31 Jul 2026 04:03:56 UTC (615 KB) Ancillary-file links: Ancillary files (details): Current browse context: cs.CL Change to browse by: References & Citation...

What's new

arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.

Key details

  • Yet the same verdicts can support wildly varying agreement numbers, depending on seemingly minor choices: the judgment scale, the retained cases, the handling of abstentions and invalid outputs, and the pooling of verdicts across items and rubric criteria.
  • The statistics that settle these choices are established, but in psychometrics, econometrics, and corpus annotation rather than in the evaluation practice that needs them.
  • We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single source-attributed analysis, and apply it to three published LLM-judge evaluations.
  • For non-degenerate binary verdicts, Pearson's $r$, Spearman's $\rho$, Kendall's $\tau_b$, the phi coefficient, and the Matthews correlation coefficient are exactly the same statistic, so reporting several repeats one number under different names.

Results & evidence

  • arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.
  • Cohen's $\kappa$ differs from them only through a marginal-mismatch factor in $(0,1]$ and shares their asymptotic variance when judge and human assign the positive verdict equally often.
  • On a rubric benchmark carrying per-criterion human labels, protocol choice alone moves reported accuracy from $0.551$ to $0.899$ and carries $\kappa$ across zero, without altering a single verdict.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

AirLLM 70B inference with single 4GB GPU

Signal 8.6 Novelty 4.0 Impact 4.8 Confidence 7.5 Actionability 3.5

Summary: Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card โ€”.

  • What happened: You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T) โ€” the largest open-source model released to date โ€” on under 4GB, because sparse.
  • Why it matters: Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card โ€” without quantization, distillation, or pruning.

What's new

Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card โ€” without quantization, distillation, or pruning.

Key details

  • You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T) โ€” the largest open-source model released to date โ€” on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.
  • - Best AI Game Sprite Generator - Best AI Facial Expression Editor - Bloome โ€” build & run AI agent teams in the cloud, zero setup [2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on...
  • Per-expert streaming loads only the experts a token actually routes to.
  • K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformer...

Results & evidence

  • You can even run 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T) โ€” the largest open-source model released to date โ€” on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.
  • - Best AI Game Sprite Generator - Best AI Facial Expression Editor - Bloome โ€” build & run AI agent teams in the cloud, zero setup [2026/07] Kimi K3 (2.8T) support: the largest open-source model runs on a single card in 3.72GB of VRAM, measured end to end on...
  • K3 brings three requirements of its own: pip install compressed-tensors flash-attn (its model code mandates flash attention regardless of what you request), a CUDA 12 build of torch, since no prebuilt flash-attn wheel exists for CUDA 13 yet, and transformer...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically

Signal 10.0 Novelty 5.1 Impact 7.8 Confidence 7.0 Actionability 6.5

Summary: AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other.

  • What happened: AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping.
  • Why it matters: It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Instead, you are programming the program.md Markdown files that provide context to the AI agents and set up your autonomous research org.

What's new

AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ri...

Key details

  • Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies.
  • The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension.
  • This repo is the story of how it all began.
  • The idea: give an AI agent a small but real LLM training setup and let it experiment autonomously overnight.

Results & evidence

  • The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension.
  • It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Reality Check

~1 min
  • nexu-io/open-design: ๐ŸŽจ The open-source Claude Design alternative. ๐Ÿ–ฅ๏ธ Local-first desktop app. ๐Ÿ–ผ๏ธ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video โ€” real files, HTML/PDF/PPTX/MP4 export. ๐Ÿค– Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
  • Primary source: yes
  • Demo available: yes
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • Show HN: Nightcrawler โ€“ A local AI pentesting agent running on a smartphone
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • AirLLM 70B inference with single 4GB GPU
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.

Lab Notes

~1 min
  • Tool/Repo of the day: nexu-io/open-design: ๐ŸŽจ The open-source Claude Design alternative. ๐Ÿ–ฅ๏ธ Local-first desktop app. ๐Ÿ–ผ๏ธ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video โ€” real files, HTML/PDF/PPTX/MP4 export. ๐Ÿค– Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK. (https://github.com/nexu-io/open-design)
  • Prompt/Workflow of the day: summarize claim -> evidence -> risk in three passes before acting.
  • Tiny snippet: `uv run python -m msd.run --scheduled`

Research Radar

~6 min

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.

  • What happened: We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single.
  • Why it matters: Under exclusion, accuracy over all cases is pinned down only to a worst-case interval as wide as the uncovered fraction.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Submission history From: Delip Rao [view email] [v1] Mon, 25 May 2026 07:31:44 UTC (269 KB) [v2] Fri, 31 Jul 2026 04:03:56 UTC (615 KB) Ancillary-file links: Ancillary files (details): Current browse context: cs.CL Change to browse by: References & Citation...

What's new

arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.

Key details

  • Yet the same verdicts can support wildly varying agreement numbers, depending on seemingly minor choices: the judgment scale, the retained cases, the handling of abstentions and invalid outputs, and the pooling of verdicts across items and rubric criteria.
  • The statistics that settle these choices are established, but in psychometrics, econometrics, and corpus annotation rather than in the evaluation practice that needs them.
  • We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single source-attributed analysis, and apply it to three published LLM-judge evaluations.
  • For non-degenerate binary verdicts, Pearson's $r$, Spearman's $\rho$, Kendall's $\tau_b$, the phi coefficient, and the Matthews correlation coefficient are exactly the same statistic, so reporting several repeats one number under different names.

Results & evidence

  • arXiv:2606.00093v2 Announce Type: replace Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels.
  • Cohen's $\kappa$ differs from them only through a marginal-mismatch factor in $(0,1]$ and shares their asymptotic variance when judge and human assign the positive verdict equally often.
  • On a rubric benchmark carrying per-criterion human labels, protocol choice alone moves reported accuracy from $0.551$ to $0.899$ and carries $\kappa$ across zero, without altering a single verdict.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema.

  • What happened: arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined.
  • Why it matters: We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios.

What's new

arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with sour...

Key details

  • We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together.
  • The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios.
  • The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms.
  • We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1.

Results & evidence

  • arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with sour...
  • The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios.
  • Computer Science > Artificial Intelligence [Submitted on 31 Jul 2026] Title:ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction View PDF HTML (experimental) Abstract:Enterprise workflows increasingly rely on agents for \emph{schema-gu...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language.

  • What happened: arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from.
  • Why it matters: arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

We conduct a systematic comparison of ten state-of-the-art LLMs and six prompting strategies that differ in the contextual information supplied (e.g., code or documentation).

What's new

arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications.

Key details

  • In this study, we explore how LLMs can be harnessed to automatically translate a neutral graph representation of fluid system models into executable code for two widely adopted simulation environments: the Python library WNTR and the Modelica Standard Library.
  • We conduct a systematic comparison of ten state-of-the-art LLMs and six prompting strategies that differ in the contextual information supplied (e.g., code or documentation).
  • For each configuration we assess the generated code using a suite of software-quality metrics and we validate the functional fidelity of the resulting simulation models by reproducing benchmark fluid system scenarios.
  • Our findings offer concrete guidance for researchers and engineers seeking to integrate LLM-driven code synthesis into model-based design pipelines.

Results & evidence

  • arXiv:2607.29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications.
  • Computer Science > Machine Learning [Submitted on 31 Jul 2026] Title:Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies View PDF HTML (experimental) Abstract:Large language models (LLMs) ha...
  • Submission history From: Jan Marius Stรผrmer [view email] [v1] Fri, 31 Jul 2026 13:09:00 UTC (675 KB) References & Citations Loading...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Forecast & Watchlist

~1 min
  • Watch: agent
  • Watch: llm
  • Watch: cs.ai
  • Watch: cs.lg
  • Watch: rss
  • Watch: cs.cl
  • Watch: python
  • Watch: benchmark

Save for Later

~7 min

mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.

Signal 10.0 Novelty 5.1 Impact 8.2 Confidence 7.0 Actionability 6.5

Summary: Straight from my .agents directory.

  • What happened: Straight from my .agents directory.
  • Why it matters: Straight from my .agents directory.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Straight from my .agents directory.

What's new

Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.

Key details

  • My agent skills that I use every day to do real engineering - not vibe coding.
  • Developing real applications is hard.
  • Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.
  • But while doing so, they take away your control and make bugs in the process hard to resolve.

Results & evidence

  • If you want to keep up with changes to these skills, and any new ones I create, you can join ~60,000 other devs on my newsletter: Two ways in, two philosophies.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

ultraworkers/claw-code: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex โ€” developed and maintained with no human intervention.

Signal 10.0 Novelty 5.1 Impact 8.2 Confidence 7.0 Actionability 6.5

Summary: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex โ€” developed and maintained with no human intervention.

  • What happened: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex โ€” developed and maintained with no human intervention.
  • Why it matters: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex โ€” developed and maintained with no human intervention.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

For file submission/navigation questions, see Navigation and file context.

What's new

Windows users can jump to the PowerShell-first Windows install and release quickstart.

Key details

  • github.com/code-yeongyu/lazycodex github.com/Yeachan-Heo/gajae-code Join the Discords: ultraworkers discord ยท gajae-code discord Important Claw Code is not the serious production project here.
  • This repository is closer to a museum exhibit than a product pitch, a crustacean-run artifact kept alive by clawed gajaes, swept and labeled by agents, and automatically maintained according to the harnesses above.
  • As already described in the project philosophy, this is not meant to be hand-operated like a normal product repo.
  • It is an agent-managed exhibit: the harnesses plan, execute, verify, label, and preserve the artifact while the crabs keep the tank running.

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2607.29402v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems synergize retrieval mechanisms with generative language models to enhance the.

  • What happened: This approach does not introduce latency but also strengthens the alignment between queries and relevant context.
  • Why it matters: arXiv:2607.29402v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems synergize retrieval mechanisms with generative language models to enhance.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

However, bridging the style gap between user queries and relevant information in document text remains a persistent challenge in retrieval-augmented systems, often addressed by runtime solutions (e.g., Hypothetical Document Embeddings (HyDE)) that attempt t...

What's new

To address these challenges, we propose Hypothetical Prompt Embeddings (HyPE), a framework that shifts the generation of hypothetical content from query time to the indexing phase.

Key details

  • However, bridging the style gap between user queries and relevant information in document text remains a persistent challenge in retrieval-augmented systems, often addressed by runtime solutions (e.g., Hypothetical Document Embeddings (HyDE)) that attempt t...
  • To address these challenges, we propose Hypothetical Prompt Embeddings (HyPE), a framework that shifts the generation of hypothetical content from query time to the indexing phase.
  • By precomputing multiple hypothetical prompts for each data chunk and embedding the chunk in place of the prompt, HyPE transforms retrieval into a question-question matching task, bypassing the need for runtime synthetic answer generation.
  • This approach does not introduce latency but also strengthens the alignment between queries and relevant context.

Results & evidence

  • arXiv:2607.29402v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems synergize retrieval mechanisms with generative language models to enhance the accuracy and relevance of responses.
  • Our experimental results on six common datasets show that HyPE can improve retrieval context precision by up to 42 percentage points and claim recall by up to 45 percentage points, compared to standard approaches, while remaining compatible with re-ranking,...

Limitations / unknowns

  • However, bridging the style gap between user queries and relevant information in document text remains a persistent challenge in retrieval-augmented systems, often addressed by runtime solutions (e.g., Hypothetical Document Embeddings (HyDE)) that attempt t...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

PISIGuard: Protect your personal and sensitive info when you chat with AI

Signal 8.4 Novelty 4.0 Impact 4.1 Confidence 7.5 Actionability 3.5

Summary: PISIGuard: Protect your personal and sensitive info when you chat with AI

  • What happened: PISIGuard: Protect your personal and sensitive info when you chat with AI
  • Why it matters: Could materially affect near-term AI workflows.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

PISIGuard: Protect your personal and sensitive info when you chat with AI

What's new

PISIGuard: Protect your personal and sensitive info when you chat with AI

Key details

  • PISIGuard: Protect your personal and sensitive info when you chat with AI

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Runthru โ€“ open-source Interactive Demos

Signal 8.4 Novelty 5.1 Impact 2.9 Confidence 7.5 Actionability 3.5

Summary: I was looking for a tool to create interactive demos / walkthroughs for web based software.

  • What happened: I was looking for a tool to create interactive demos / walkthroughs for web based software.
  • Why it matters: I was looking for a tool to create interactive demos / walkthroughs for web based software.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

I was looking for a tool to create interactive demos / walkthroughs for web based software.

What's new

I was looking for a tool to create interactive demos / walkthroughs for web based software.

Key details

  • I was amazed that many companies charge well over $1000 per month for something pretty basic.

    I made this which so far works well for me.

  • You can point it at a site and either let AI (byo OpenAI or Anthropic keys) record and polish the whole thing or do it manually yourself (or a mixture of both).

    Here is an example interactive demo:

    ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

    Signal 7.3 Novelty 6.2 Impact 2.0 Confidence 3.8 Actionability 3.5

    Summary: ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

    • What happened: ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
    • Why it matters: Could materially affect near-term AI workflows.
    • What to do: Track for corroboration and benchmark data before adopting.
    Deep

    Context

    ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

    What's new

    ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

    Key details

    • ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

    Results & evidence

    • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

    Limitations / unknowns

    • Generalization outside curated tasks is still unclear.

    Next-step validation checks

    • Reproduce one claim with a public baseline and fixed evaluation settings.
    • Check robustness on out-of-distribution or long-context cases.
    • Track whether independent teams report matching results.