Morning Singularity Digest - 2026-09-26

Estimated total read • ~32 min

Skim fast, dive deep only where it matters.

2-minute skim 10-minute read Deep dive optional
Contents

Front Page

~7 min

nexu-io/open-design: 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK.

Signal 10.0 Novelty 7.3 Impact 7.8 Confidence 7.0 Actionability 6.5

Summary: 🎨 Best DeepSeek Harness Design Plugin.

  • What happened: 🎨 Best DeepSeek Harness Design Plugin.
  • Why it matters: 🎨 Best DeepSeek Harness Design Plugin.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

🎨 Best DeepSeek Harness Design Plugin.

What's new

🖥️ Local-first native desktop app for macOS and Windows.

Key details

  • The open-source Claude Design alternative.
  • 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export.
  • 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK.
  • ⚡ OpenDesign Cloud — the official model service.

Results & evidence

  • 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK.
  • One recharge to use both agent and image models inside OpenDesign: GPT, Claude, and DeepSeek for agents; GPT Image 2.0, Seedream 5.0 Pro, and Nano Banana 2.0 for images.

Limitations / unknowns

  • OpenDesign members can use both models without limits for two weeks, directly inside the app.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.

Signal 10.0 Novelty 5.1 Impact 8.4 Confidence 7.0 Actionability 6.5

Summary: Straight from my .agents directory.

  • What happened: Straight from my .agents directory.
  • Why it matters: Straight from my .agents directory.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Straight from my .agents directory.

What's new

Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.

Key details

  • My agent skills that I use every day to do real engineering - not vibe coding.
  • Developing real applications is hard.
  • Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process.
  • But while doing so, they take away your control and make bugs in the process hard to resolve.

Results & evidence

  • If you want to keep up with changes to these skills, and any new ones I create, you can join ~60,000 other devs on my newsletter: Two ways in, two philosophies.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning.

  • What happened: While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask.
  • Why it matters: arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles.

What's new

In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.

Key details

  • While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the t...
  • In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.
  • Computer Science > Computation and Language [Submitted on 24 Sep 2026] Title:TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) View PDF HTML (experimental) Abstract:Arabic-language stance detection re...
  • Bibliographic and Citation Tools Bibliographic Explorer (What is the Explorer?) Connected Papers (What is Connected Papers?) Litmaps (What is Litmaps?) scite Smart Citations (What are Smart Citations?) Code, Data and Media Associated with this Article alpha...

Results & evidence

  • arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles.
  • Computer Science > Computation and Language [Submitted on 24 Sep 2026] Title:TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) View PDF HTML (experimental) Abstract:Arabic-language stance detection re...

Limitations / unknowns

  • While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the t...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.

  • What happened: arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.
  • Why it matters: Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.

What's new

arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.

Key details

  • We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted.
  • The measured phenomenon is unstable to begin with.
  • Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect.
  • Auditing the evaluation weakens its conclusions further, and this is our main contribution.

Results & evidence

  • arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.
  • We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted.
  • Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Better prompt caching for GPT-6

Signal 7.3 Novelty 4.0 Impact 2.0 Confidence 3.0 Actionability 5.2

Summary: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

  • What happened: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
  • Why it matters: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

What's new

Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Key details

  • Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Results & evidence

  • Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

What Changed Overnight

~1 min
  • New: JuliusBrussee/caveman: 🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
  • New: stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
  • New: tt-a1i/archify: Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.
  • New: One Month Without AI
  • New: Understanding the Impact of LLM Watermarking on AI Agent Behavior
  • New: CEO of Mistral: AI is software. It can be controlled
  • Removed: ultraworkers/claw-code: An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention. (fell below rank threshold)
  • Removed: VoltAgent/awesome-design-md: A collection of DESIGN.md files analysis by popular brand design systems. Drop one into your project and let coding agents generate a matching UI. (fell below rank threshold)
  • Removed: multica-ai/andrej-karpathy-skills: A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls. (fell below rank threshold)
  • Removed: Pistis Technical Report (fell below rank threshold)
  • What to do now:
  • Validate with one small internal benchmark and compare against your current baseline this week.
  • Track for corroboration and benchmark data before adopting.

Deep Dives

~6 min

TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning.

  • What happened: While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask.
  • Why it matters: arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles.

What's new

In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.

Key details

  • While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the t...
  • In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.
  • Computer Science > Computation and Language [Submitted on 24 Sep 2026] Title:TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) View PDF HTML (experimental) Abstract:Arabic-language stance detection re...
  • Bibliographic and Citation Tools Bibliographic Explorer (What is the Explorer?) Connected Papers (What is Connected Papers?) Litmaps (What is Litmaps?) scite Smart Citations (What are Smart Citations?) Code, Data and Media Associated with this Article alpha...

Results & evidence

  • arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles.
  • Computer Science > Computation and Language [Submitted on 24 Sep 2026] Title:TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) View PDF HTML (experimental) Abstract:Arabic-language stance detection re...

Limitations / unknowns

  • While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the t...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

One Month Without AI

Signal 8.9 Novelty 4.0 Impact 5.9 Confidence 6.2 Actionability 3.5

Summary: Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner.

  • What happened: Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner.
  • Why it matters: Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner.

What's new

Several months ago, I decided that AI contributions were no longer welcome in a FOSS project I am building and maintaining - LibreWeddingPlanner.

Key details

  • It’s not that it got a lot of contributions with AI — actually all contributions I’ve had are translations and feature requests — but I wanted to avoid future drama and have a position against AI.
  • However, although I was not using AI for my FOSS contributions, I kept using it at work.
  • In my workplace, as well as in many of my developer friends’, using AI to work is very, very common.
  • Not without a degree of shame, let me tell you about this experience, and how it was turning me dumber, lazy, and a worse developer.

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • However, although I was not using AI for my FOSS contributions, I kept using it at work.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically

Signal 10.0 Novelty 5.1 Impact 7.8 Confidence 7.0 Actionability 6.5

Summary: AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other.

  • What happened: AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping.
  • Why it matters: It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Instead, you are programming the program.md Markdown files that provide context to the AI agents and set up your autonomous research org.

What's new

AI agents running research on single-GPU nanochat training automatically One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ri...

Key details

  • Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies.
  • The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension.
  • This repo is the story of how it all began.
  • The idea: give an AI agent a small but real LLM training setup and let it experiment autonomously overnight.

Results & evidence

  • The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension.
  • It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Reality Check

~1 min
  • nexu-io/open-design: 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK.
  • Primary source: yes
  • Demo available: yes
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • mattpocock/skills: Skills for Real Engineers. Straight from my .agents directory.
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: no
  • Baselines/ablations: no
  • Third-party corroboration: no
  • Reproducibility details: yes
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: yes
  • Baselines/ablations: yes
  • Third-party corroboration: no
  • Reproducibility details: no
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.
  • How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
  • Primary source: yes
  • Demo available: no
  • Benchmarks/evals: yes
  • Baselines/ablations: yes
  • Third-party corroboration: no
  • Reproducibility details: no
  • What would change my mind:
  • Independent replication with comparable or better results.
  • Public benchmark numbers with clear baseline comparisons.
  • Likely failure mode: Performance may collapse outside curated demos or narrow tasks.

Lab Notes

~1 min
  • Tool/Repo of the day: nexu-io/open-design: 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK. (https://github.com/nexu-io/open-design)
  • Prompt/Workflow of the day: summarize claim -> evidence -> risk in three passes before acting.
  • Tiny snippet: `uv run python -m msd.run --scheduled`

Research Radar

~6 min

TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning.

  • What happened: While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask.
  • Why it matters: arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles.

What's new

In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.

Key details

  • While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the t...
  • In this approach, the target, predicted sentiment, and text are combined into a single prompt whose $\texttt{[MASK]}$ prediction is restricted to a verbalizer-constrained label vocabulary.
  • Computer Science > Computation and Language [Submitted on 24 Sep 2026] Title:TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) View PDF HTML (experimental) Abstract:Arabic-language stance detection re...
  • Bibliographic and Citation Tools Bibliographic Explorer (What is the Explorer?) Connected Papers (What is Connected Papers?) Litmaps (What is Litmaps?) scite Smart Citations (What are Smart Citations?) Code, Data and Media Associated with this Article alpha...

Results & evidence

  • arXiv:2609.29733v1 Announce Type: cross Abstract: Arabic-language stance detection remains challenging, and previous shared-task systems have largely relied on multitask learning and ensembles.
  • Computer Science > Computation and Language [Submitted on 24 Sep 2026] Title:TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar) View PDF HTML (experimental) Abstract:Arabic-language stance detection re...

Limitations / unknowns

  • While these systems achieve state-of-the-art performance, their applicability and transferability are limited by the additional complexity introduced by multitask learning.To reduce this complexity, we introduce $\texttt{CLASP-Ar}$, which reformulates the t...

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 8.3 Actionability 5.2

Summary: arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.

  • What happened: arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.
  • Why it matters: Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.

What's new

arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.

Key details

  • We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted.
  • The measured phenomenon is unstable to begin with.
  • Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect.
  • Auditing the evaluation weakens its conclusions further, and this is our main contribution.

Results & evidence

  • arXiv:2609.30074v1 Announce Type: cross Abstract: Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table.
  • We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted.
  • Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

FB-GDM: Fully-Bayesian Guided Diffusion Models for High-Dimensional Linear Inverse Problems via Unsupervised Variational Inference

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 7.5 Actionability 5.2

Summary: arXiv:2609.29216v1 Announce Type: new Abstract: Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling.

  • What happened: We introduce FB-GDM, a fully-Bayesian guided diffusion method that removes this calibration step.
  • Why it matters: arXiv:2609.29216v1 Announce Type: new Abstract: Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.29216v1 Announce Type: new Abstract: Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling (DPS) and Pseudoinverse-Guided Diffusion Models ($\Pi$GDM), rely on scalar hyp...

What's new

arXiv:2609.29216v1 Announce Type: new Abstract: Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling (DPS) and Pseudoinverse-Guided Diffusion Models ($\Pi$GDM), rely on scalar hyp...

Key details

  • We introduce FB-GDM, a fully-Bayesian guided diffusion method that removes this calibration step.
  • Starting from the Gaussian approximation of $\Pi$GDM, we derive a closed-form conditional score that depends on two precision parameters (inverse variances), one associated with the denoising approximation and one with the observation likelihood, and treat...
  • A separable factorization makes each update scale linearly with the number of pixels, so the inference stays tractable at full image resolution, at a cost comparable to one $\Pi$GDM run.
  • FB-GDM requires neither the noise level nor the ground truth: its only inputs are the observation and the forward operator.

Results & evidence

  • arXiv:2609.29216v1 Announce Type: new Abstract: Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling (DPS) and Pseudoinverse-Guided Diffusion Models ($\Pi$GDM), rely on scalar hyp...
  • (i) The precision parameters, inferred from the observation alone, allow FB-GDM to outperform $\Pi$GDM at its nominal setting, even when the latter is given the true noise level, by up to 14 dB depending on the operator, and to match the ground-truth-calibr...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Forecast & Watchlist

~1 min
  • Watch: cs.ai
  • Watch: cs.lg
  • Watch: rss
  • Watch: cs.cl
  • Watch: python
  • Watch: benchmark
  • Watch: eval
  • Watch: repo

Save for Later

~9 min

DietrichGebert/ponytail: Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

Signal 10.0 Novelty 5.1 Impact 8.1 Confidence 7.0 Actionability 6.5

Summary: Makes your AI agent think like the laziest senior dev in the room.

  • What happened: Makes your AI agent think like the laziest senior dev in the room.
  • Why it matters: ~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe Measured on real Claude Code sessions editing a real open-source repo (FastAPI + React), against the.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

Makes your AI agent think like the laziest senior dev in the room.

What's new

Makes your AI agent think like the laziest senior dev in the room.

Key details

  • The best code is the code you never wrote.
  • ~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe Measured on real Claude Code sessions editing a real open-source repo (FastAPI + React), against the same agent with no skill.
  • ~54% is the mean across 12 feature tasks (Haiku 4.5, n=4); it reaches 94% where an agent over-builds (a date picker) and is near zero where the code is already minimal.
  • ponytail keeps every safety guard while a bare "write one-liners" prompt drops one.

Results & evidence

  • ~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe Measured on real Claude Code sessions editing a real open-source repo (FastAPI + React), against the same agent with no skill.
  • ~54% is the mean across 12 feature tasks (Haiku 4.5, n=4); it reaches 94% where an agent over-builds (a date picker) and is near zero where the code is already minimal.
  • (The earlier single-shot benchmark reported 80-94% as a flat figure; against a fair agentic baseline that is the per-task ceiling, not the average.) Full writeup · reproduce it.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

JuliusBrussee/caveman: 🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.

Signal 10.0 Novelty 5.1 Impact 7.9 Confidence 7.0 Actionability 6.5

Summary: 🪨 why use many token when few token do trick.

  • What happened: 🪨 why use many token when few token do trick.
  • Why it matters: 🪨 why use many token when few token do trick.
  • What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep

Context

🪨 why use many token when few token do trick.

What's new

🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...

Key details

  • Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
  • Your AI coding agent bills by the word and writes like it knows that.
  • 🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
  • npx skills add JuliusBrussee/caveman -g → Quick Start See it · Quick Start · The Numbers · How it compares · In the Wild · The Skill · The Proxy · Wrap · Your own app · When to Skip · Docs | 🗣️ Normal agent · 69 tokens | Caveman agent · 19 tokens | |---|---...

Results & evidence

  • Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
  • 🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
  • npx skills add JuliusBrussee/caveman -g → Quick Start See it · Quick Start · The Numbers · How it compares · In the Wild · The Skill · The Proxy · Wrap · Your own app · When to Skip · Docs | 🗣️ Normal agent · 69 tokens | Caveman agent · 19 tokens | |---|---...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

Signal 9.4 Novelty 4.0 Impact 2.0 Confidence 7.5 Actionability 5.2

Summary: arXiv:2609.29505v1 Announce Type: new Abstract: Diffusion inference repeatedly evaluates the same large network.

  • What happened: arXiv:2609.29505v1 Announce Type: new Abstract: Diffusion inference repeatedly evaluates the same large network.
  • Why it matters: arXiv:2609.29505v1 Announce Type: new Abstract: Diffusion inference repeatedly evaluates the same large network.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

arXiv:2609.29505v1 Announce Type: new Abstract: Diffusion inference repeatedly evaluates the same large network.

What's new

arXiv:2609.29505v1 Announce Type: new Abstract: Diffusion inference repeatedly evaluates the same large network.

Key details

  • We ask whether pretrained weights alone can identify residual branches that need not be recomputed throughout the trajectory.
  • Our \textbf{Spectral Concentration Ratio (SCR)} measures leading-versus-tail singular-value energy.
  • Combined with Frobenius magnitude, it yields an offline sensitivity proxy and a deterministic lifetime for each scheduled unit.
  • A frozen unit reuses its cached residual-branch update while the current residual stream and all external conditioning continue to propagate.

Results & evidence

  • arXiv:2609.29505v1 Announce Type: new Abstract: Diffusion inference repeatedly evaluates the same large network.
  • At matched layer-step budgets, SCR/Frobenius preserves quality better than random, depth, norm, stable-rank, and Frobenius--stable-rank schedules on LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL.
  • The complete captured-graph system reaches $2.8\times$--$3.0\times$ wall-clock speedup over eager inference.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents

Signal 8.4 Novelty 5.1 Impact 2.7 Confidence 6.2 Actionability 5.2

Summary: Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents

  • What happened: Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents
  • Why it matters: Could materially affect near-term AI workflows.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents

What's new

Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents

Key details

  • Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Focal Prompt: tools for studying how AI systems allocate attention

Signal 8.4 Novelty 4.0 Impact 2.6 Confidence 6.2 Actionability 5.2

Summary: Focal Prompt: tools for studying how AI systems allocate attention

  • What happened: Focal Prompt: tools for studying how AI systems allocate attention
  • Why it matters: Could materially affect near-term AI workflows.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

Focal Prompt: tools for studying how AI systems allocate attention

What's new

Focal Prompt: tools for studying how AI systems allocate attention

Key details

  • Focal Prompt: tools for studying how AI systems allocate attention

Results & evidence

  • No hard numbers surfaced in the source text; treat claims as directional until benchmarks appear.

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.

Understanding the Impact of LLM Watermarking on AI Agent Behavior

Signal 8.6 Novelty 5.1 Impact 5.1 Confidence 6.2 Actionability 3.5

Summary: The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, Anthropic announced that future Claude models would embed an invisible watermark in.

  • What happened: The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, Anthropic announced that future Claude models would embed an invisible.
  • Why it matters: The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, Anthropic announced that future Claude models would embed an invisible.
  • What to do: Track for corroboration and benchmark data before adopting.
Deep

Context

The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, Anthropic announced that future Claude models would embed an invisible watermark in their output [1], [2], and subsequently disclosed that the watermark is based...

What's new

Text watermarking itself is not new, but its deployment now has regulatory relevance.

Key details

  • Text watermarking itself is not new, but its deployment now has regulatory relevance.
  • Article 50(2) of the EU AI Act [4] requires providers of AI systems generating synthetic text to mark their outputs in a machine-readable format and make them detectable as artificially generated or manipulated, using technical solutions that are effective,...
  • Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token.
  • At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection.

Results & evidence

  • The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Recently, Anthropic announced that future Claude models would embed an invisible watermark in their output [1], [2], and subsequently disclosed that the watermark is based...
  • Article 50(2) of the EU AI Act [4] requires providers of AI systems generating synthetic text to mark their outputs in a machine-readable format and make them detectable as artificially generated or manipulated, using technical solutions that are effective,...

Limitations / unknowns

  • Generalization outside curated tasks is still unclear.

Next-step validation checks

  • Reproduce one claim with a public baseline and fixed evaluation settings.
  • Check robustness on out-of-distribution or long-context cases.
  • Track whether independent teams report matching results.