Morning Singularity Digest - 2026-09-30

Estimated total read • ~35 min

Skim fast, dive deep only where it matters.

2-minute skim 10-minute read Deep dive optional
Contents

Front Page

~8 min

career-ops-hq/career-ops: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)

Signal 10.0 Novelty 6.2 Impact 7.7 Confidence 7.8 Actionability 6.5

Summary: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in.

  • What happened: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs.
  • Why it matters: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) The open-source A...

What's new

Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) The open-source A...

Key details

  • English | Español | Deutsch | Français | Português (Brasil) | 한국어 | 日本語 | 简体中文 | 繁體中文 | Українська | Русский | Polski | Dansk | தமிழ் | العربية | हिन्दी | Türkçe Months of sending CVs into silence.
  • On your machine, it tells you if it's still open and if it fits.
  • It tailors your CV and drafts your answers.
  • My own search, partway through, with the UI in Spanish.

Results & evidence

  • Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) The open-source A...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.

Signal 10.0 Novelty 5.1 Impact 8.4 Confidence 7.0 Actionability 6.5

Summary: Straight from my .agents directory.

  • What happened: Straight from my .agents directory.
  • Why it matters: Straight from my .agents directory.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Straight from my .agents directory.

What's new

Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.

Key details

  • My agent skills that I use every day to do real engineering - not vibe coding.
  • Developing real applications is hard.
  • Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.
  • But while doing so, they take away your control and make bugs in the process hard to resolve.

Results & evidence

  • If you want to keep up with changes to these skills, and any new ones I create, you can join ~60,000 other devs on my newsletter: Two ways in, two philosophies.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

Signal 9.4 Novelty 6.2 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios.

  • What happened: To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering.
  • Why it matters: arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE).
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.

What's new

To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability.

Key details

  • However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation.
  • To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages.
  • Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios.
  • To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability.

Results & evidence

  • arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.
  • To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages.
  • Computer Science > Software Engineering [Submitted on 5 Aug 2026 (v1), last revised 29 Sep 2026 (this version, v2)] Title:Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports View PDF HTML (experimental) Abstract:Coding agen...

Limitations / unknowns

  • However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation.
  • Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potentia...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as.

  • What happened: arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files.
  • Why it matters: Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md.

What's new

arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md.

Key details

  • Although this practice is strongly encouraged by agent developers, there is currently no rigorous investigation into whether such context files are actually effective for real-world tasks.
  • In this work, we study this question and evaluate coding agents' task completion performance in two complementary settings: established SWE-bench tasks from popular repositories, with LLM-generated context files, and a novel collection of issues from reposi...
  • Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average.
  • This observation holds across different LLMs, coding agents, and for both LLM-generated and developer-committed context files.

Results & evidence

  • arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md.
  • Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average.
  • Computer Science > Software Engineering [Submitted on 12 Feb 2026 (v1), last revised 29 Sep 2026 (this version, v3)] Title:Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Glance – Build, evaluate, and monitor AI agents

Signal 8.4 Novelty 5.1 Impact 2.6 Confidence 8.2 Actionability 3.5

Summary: Show HN: Glance – Build, evaluate, and monitor AI agents

  • What happened: Show HN: Glance – Build, evaluate, and monitor AI agents
  • Why it matters: Could materially affect near-term AI workflows.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Show HN: Glance – Build, evaluate, and monitor AI agents

What's new

Show HN: Glance – Build, evaluate, and monitor AI agents

Key details

  • Show HN: Glance – Build, evaluate, and monitor AI agents

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

What Changed Overnight

~1 min
  • New: career-ops-hq/career-ops: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)
  • New: JuliusBrussee/caveman: 🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
  • New: karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically
  • New: Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
  • New: The AI Race Just Got Awkward
  • New: Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
  • Removed: affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. (fell below rank threshold)
  • Removed: VoltAgent/awesome-design-md: A collection of DESIGN.md files analysis by popular brand design systems. Drop one into your project and let coding agents generate a matching UI. (fell below rank threshold)
  • Removed: multica-ai/andrej-karpathy-skills: A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls. (fell below rank threshold)
  • Removed: Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models (fell below rank threshold)
  • What to do now:
  • Validate with one small internal benchmark and compare against your current baseline this week.
  • Track for corroboration and benchmark data before adopting.

Deep Dives

~7 min

career-ops-hq/career-ops: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…)

Signal 10.0 Novelty 6.2 Impact 7.7 Confidence 7.8 Actionability 6.5

Summary: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in.

  • What happened: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs.
  • Why it matters: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) The open-source A...

What's new

Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) The open-source A...

Key details

  • English | Español | Deutsch | Français | Português (Brasil) | 한국어 | 日本語 | 简体中文 | 繁體中文 | Українська | Русский | Polski | Dansk | தமிழ் | العربية | हिन्दी | Türkçe Months of sending CVs into silence.
  • On your machine, it tells you if it's still open and if it fits.
  • It tailors your CV and drafts your answers.
  • My own search, partway through, with the UI in Spanish.

Results & evidence

  • Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) The open-source A...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

Signal 9.4 Novelty 6.2 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios.

  • What happened: To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering.
  • Why it matters: arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE).
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.

What's new

To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability.

Key details

  • However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation.
  • To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages.
  • Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios.
  • To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability.

Results & evidence

  • arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.
  • To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages.
  • Computer Science > Software Engineering [Submitted on 5 Aug 2026 (v1), last revised 29 Sep 2026 (this version, v2)] Title:Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports View PDF HTML (experimental) Abstract:Coding agen...

Limitations / unknowns

  • However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation.
  • Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potentia...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

The AI Race Just Got Awkward

Signal 9.5 Novelty 4.0 Impact 6.4 Confidence 6.2 Actionability 3.5

Summary: If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.

  • What happened: If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.
  • Why it matters: If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

It is a mind-blowing optimization that basically dropped the KV cache footprint for certain use cases that use a long session context, like coding, by a factor of roughly 437x compared with DeepSeek-V1.

What's new

If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.

Key details

  • The Distillation Drama Not a week goes by when Anthropic doesnt release another article on how the Chinese are distilling their models, becoming a danger to humanity itself, etc.
  • Its beneficial for them to say that because it sets the ground for these models to be restrained legally and regulatorily later on.
  • But its clear that the days of mindless distilling are over.
  • The new game in town is adopting Chinese labs advances.

Results & evidence

  • It is a mind-blowing optimization that basically dropped the KV cache footprint for certain use cases that use a long session context, like coding, by a factor of roughly 437x compared with DeepSeek-V1.
  • They were the first ones to release the MLA architecture, which compressed the cache by roughly 15x, and then followed it up with Compressed Sparse Attention and Heavily Compressed Attention.
  • The latest DeepSeek-V4.1-Flash pushes it even further with CSA2, cross-layer cache reuse, a causal encoder-decoder architecture, and FP4 caching, bringing the global KV cache down to 890 bytes per token.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Reality Check

~1 min
  • mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • The AI Race Just Got Awkward
  • Primary source: no
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: no
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.

Lab Notes

~1 min
  • Tool/Repo of the day: career-ops-hq/career-ops: Open-source AI job search agent: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Code, Codex, OpenCode, Antigravity…) (https://github.com/career-ops-hq/career-ops)
  • Prompt/Workflow of the day: summarize claim -> evidence -> risk in three passes before acting.
  • Tiny snippet: `uv run python -m msd.run --scheduled`

Research Radar

~6 min

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

Signal 9.4 Novelty 6.2 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios.

  • What happened: To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering.
  • Why it matters: arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE).
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.

What's new

To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability.

Key details

  • However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation.
  • To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages.
  • Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios.
  • To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability.

Results & evidence

  • arXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase.
  • To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages.
  • Computer Science > Software Engineering [Submitted on 5 Aug 2026 (v1), last revised 29 Sep 2026 (this version, v2)] Title:Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports View PDF HTML (experimental) Abstract:Coding agen...

Limitations / unknowns

  • However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation.
  • Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potentia...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 9.5 Actionability 6.5

Summary: arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as.

  • What happened: arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files.
  • Why it matters: Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md.

What's new

arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md.

Key details

  • Although this practice is strongly encouraged by agent developers, there is currently no rigorous investigation into whether such context files are actually effective for real-world tasks.
  • In this work, we study this question and evaluate coding agents' task completion performance in two complementary settings: established SWE-bench tasks from popular repositories, with LLM-generated context files, and a novel collection of issues from reposi...
  • Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average.
  • This observation holds across different LLMs, coding agents, and for both LLM-generated and developer-committed context files.

Results & evidence

  • arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md.
  • Surprisingly, we find that providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average.
  • Computer Science > Software Engineering [Submitted on 12 Feb 2026 (v1), last revised 29 Sep 2026 (this version, v3)] Title:Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems

Signal 9.4 Novelty 5.1 Impact 2.0 Confidence 8.7 Actionability 6.5

Summary: arXiv:2609.32490v2 Announce Type: replace Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task.

  • What happened: To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and.
  • Why it matters: arXiv:2609.32490v2 Announce Type: replace Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

We refer to such problems as progressively specified tasks.

What's new

We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management.

Key details

  • In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution.
  • We refer to such problems as progressively specified tasks.
  • To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence.
  • We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management.

Results & evidence

  • arXiv:2609.32490v2 Announce Type: replace Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution.
  • Computer Science > Artificial Intelligence [Submitted on 26 Sep 2026 (v1), last revised 29 Sep 2026 (this version, v2)] Title:RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems View PDF HTML (experimental) Abstract:LLM-base...
  • Submission history From: Yuchen Song [view email] [v1] Sat, 26 Sep 2026 11:33:53 UTC (431 KB) [v2] Tue, 29 Sep 2026 08:06:13 UTC (431 KB) References & Citations Loading...

Limitations / unknowns

  • RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Forecast & Watchlist

~1 min
  • Watch: cs.ai
  • Watch: cs.lg
  • Watch: rss
  • Watch: cs.cl
  • Watch: python
  • Watch: benchmark
  • Watch: eval
  • Watch: repo

Save for Later

~10 min

JuliusBrussee/caveman: 🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.

Signal 10.0 Novelty 5.1 Impact 7.9 Confidence 7.0 Actionability 6.5

Summary: 🪨 why use many token when few token do trick.

  • What happened: 🪨 why use many token when few token do trick.
  • Why it matters: 🪨 why use many token when few token do trick.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

🪨 why use many token when few token do trick.

What's new

🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...

Key details

  • Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
  • Your AI coding agent bills by the word and writes like it knows that.
  • 🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
  • npx skills add JuliusBrussee/caveman -g → Quick Start See it · Quick Start · The Numbers · How it compares · In the Wild · The Skill · The Proxy · Wrap · Your own app · When to Skip · Docs | 🗣️ Normal agent · 69 tokens | Caveman agent · 19 tokens | |---|---...

Results & evidence

  • Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
  • 🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
  • npx skills add JuliusBrussee/caveman -g → Quick Start See it · Quick Start · The Numbers · How it compares · In the Wild · The Skill · The Proxy · Wrap · Your own app · When to Skip · Docs | 🗣️ Normal agent · 69 tokens | Caveman agent · 19 tokens | |---|---...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

addyosmani/agent-skills: Production-grade engineering skills for AI coding agents.

Signal 10.0 Novelty 5.1 Impact 7.8 Confidence 7.0 Actionability 6.5

Summary: Production-grade engineering skills for AI coding agents.

  • What happened: Production-grade engineering skills for AI coding agents.
  • Why it matters: Production-grade engineering skills for AI coding agents.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Production-grade engineering skills for AI coding agents.

What's new

Production-grade engineering skills for AI coding agents.

Key details

  • Skills encode the workflows, quality gates, and best practices that senior engineers use when building software.
  • These ones are packaged so AI agents follow them consistently across every phase of development.
  • DEFINE PLAN BUILD VERIFY REVIEW SHIP ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ │ Idea │ ───▶ │ Spec │ ───▶ │ Code │ ───▶ │ Test │ ───▶ │ QA │ ───▶ │ Go │ │Refine│ │ PRD │ │ Impl │ │Debug │ │ Gate │ │ Live │ └──────┘ └──────┘ └──────┘ └──────┘ └─...
  • Each one activates the right skills automatically.

Results & evidence

  • The open skills CLI installs into 70+ agents (Claude Code, Cursor, Codex, Copilot, Cline, and more): npx skills add addyosmani/agent-skills # install all 25 skills npx skills add addyosmani/agent-skills --list # browse before installing Or grab individual s...

Limitations / unknowns

  • It removes the human stepping between tasks, not the verification: every task is still test-driven and committed individually, and it pauses on failures or risky steps.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

LongCat-DeepResearch Technical Report

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.7 Actionability 6.5

Summary: arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for.

  • What happened: arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent.
  • Why it matters: Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop.

What's new

arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports.

Key details

  • The workflow separates global planning from detailed investigation and coordinates revision at the section level.
  • Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec.
  • Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop.
  • Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting.

Results & evidence

  • arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports.
  • LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics.
  • On an in-house benchmark, it scores 76.04, ranking second among four compared systems.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: Paperweight – local-first, open-source email cleanup and privacy tool

Signal 8.4 Novelty 6.2 Impact 2.6 Confidence 7.5 Actionability 3.5

Summary: Hi HN,

After building for 7 months, I'm excited to announce Paperweight V1.

The idea is that your inbox knows who you ever interacted with and knows where your data lives.

  • What happened: Hi HN,

    After building for 7 months, I'm excited to announce Paperweight V1.

    The idea is that your inbox knows who you ever interacted with and knows where your.

  • Why it matters: Hi HN,

    After building for 7 months, I'm excited to announce Paperweight V1.

    The idea is that your inbox knows who you ever interacted with and knows where your.

  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Hi HN,

After building for 7 months, I'm excited to announce Paperweight V1.

The idea is that your inbox knows who you ever interacted with and knows where your data lives.

What's new

It never felt right that with a lot of alternatives, in order to reclaim your privacy you first had to hand over your data.

Key details

  • Every account you ever created, every service you signed up for, and every online purchase is connected to your email.
  • Most people have hundreds of accounts they've forgotten about, creating security risks and privacy exposure.

    Paperweight scans your email history locally to find old accounts, companies, mailing lists, data breaches, and personal information.

  • You can then unsubscribe, clean up unwanted email, send privacy/GDPR requests, and keep track of your digital footprint.

    The main constraint from the start was privacy.

  • It never felt right that with a lot of alternatives, in order to reclaim your privacy you first had to hand over your data.

Results & evidence

  • Hi HN,

    After building for 7 months, I'm excited to announce Paperweight V1.

    The idea is that your inbox knows who you ever interacted with and knows where your data lives.

  • I haven't been a customer with them for 8+ years, but they still had all my data in their systems.

    Happy to answer any questions or hear what you think.

    GitHub: http...

Limitations / unknowns

  • Most people have hundreds of accounts they've forgotten about, creating security risks and privacy exposure.

    Paperweight scans your email history locally to find old accounts, companies, mailing lists, data breaches, and personal information.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Show HN: WattzGOAT – an intentionally vulnerable web app for security training

Signal 8.4 Novelty 4.0 Impact 2.6 Confidence 7.5 Actionability 3.5

Summary: Built this as a hands-on lab for teaching web application security.

  • What happened: Built this as a hands-on lab for teaching web application security.
  • Why it matters: Built this as a hands-on lab for teaching web application security.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Built this as a hands-on lab for teaching web application security.

What's new

Built this as a hands-on lab for teaching web application security.

Key details

  • It is a fictional electricity utility's customer portal with 48 deliberately planted flags (most of OWASP Top 10) you find and exploit CTF-style, including a simulated (rule-based, not a real model) AI assistant with its own prompt-injection-style vuln...

Results & evidence

  • It is a fictional electricity utility's customer portal with 48 deliberately planted flags (most of OWASP Top 10) you find and exploit CTF-style, including a simulated (rule-based, not a real model) AI assistant with its own prompt-injection-style vuln...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Better prompt caching for GPT-6

Signal 7.3 Novelty 4.0 Impact 2.0 Confidence 3.0 Actionability 5.2

Summary: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

  • What happened: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
  • Why it matters: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

What's new

Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Key details

  • Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Results & evidence

  • Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.