Source: github | Overall 7.8/10 | Corroboration: 1
Signal 10.0
Novelty 5.1
Impact 7.9
Confidence 7.0
Actionability 6.5
Summary: 🪨 why use many token when few token do trick.
- What happened: 🪨 why use many token when few token do trick.
- Why it matters: 🪨 why use many token when few token do trick.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
🪨 why use many token when few token do trick.
What's new
🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
Key details
- Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
- Your AI coding agent bills by the word and writes like it knows that.
- 🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
- npx skills add JuliusBrussee/caveman -g → Quick Start See it · Quick Start · The Numbers · How it compares · In the Wild · The Skill · The Proxy · Wrap · Your own app · When to Skip · Docs | 🗣️ Normal agent · 69 tokens | Caveman agent · 19 tokens | |---|---...
Results & evidence
- Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
- 🏆 #1 on GitHub Trending · July 2026 · 🥇 #1 Repository of the Day on Trendshift · April 2026 #1 on Hacker News · 904 points · 366 comments · #8 Product of the Day on Product Hunt 📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style outpu...
- npx skills add JuliusBrussee/caveman -g → Quick Start See it · Quick Start · The Numbers · How it compares · In the Wild · The Skill · The Proxy · Wrap · Your own app · When to Skip · Docs | 🗣️ Normal agent · 69 tokens | Caveman agent · 19 tokens | |---|---...
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: github | Overall 7.8/10 | Corroboration: 1
Signal 10.0
Novelty 5.1
Impact 7.8
Confidence 7.0
Actionability 6.5
Summary: Production-grade engineering skills for AI coding agents.
- What happened: Production-grade engineering skills for AI coding agents.
- Why it matters: Production-grade engineering skills for AI coding agents.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
Production-grade engineering skills for AI coding agents.
What's new
Production-grade engineering skills for AI coding agents.
Key details
- Skills encode the workflows, quality gates, and best practices that senior engineers use when building software.
- These ones are packaged so AI agents follow them consistently across every phase of development.
- DEFINE PLAN BUILD VERIFY REVIEW SHIP ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ │ Idea │ ───▶ │ Spec │ ───▶ │ Code │ ───▶ │ Test │ ───▶ │ QA │ ───▶ │ Go │ │Refine│ │ PRD │ │ Impl │ │Debug │ │ Gate │ │ Live │ └──────┘ └──────┘ └──────┘ └──────┘ └─...
- Each one activates the right skills automatically.
Results & evidence
- The open skills CLI installs into 70+ agents (Claude Code, Cursor, Codex, Copilot, Cline, and more): npx skills add addyosmani/agent-skills # install all 25 skills npx skills add addyosmani/agent-skills --list # browse before installing Or grab individual s...
Limitations / unknowns
- It removes the human stepping between tasks, not the verification: every task is still test-driven and committed individually, and it pauses on failures or risky steps.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: arxiv | Overall 6.2/10 | Corroboration: 1
Signal 9.4
Novelty 4.0
Impact 2.0
Confidence 8.7
Actionability 6.5
Summary: arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for.
- What happened: arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent.
- Why it matters: Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop.
What's new
arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports.
Key details
- The workflow separates global planning from detailed investigation and coordinates revision at the section level.
- Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec.
- Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop.
- Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting.
Results & evidence
- arXiv:2609.36071v1 Announce Type: new Abstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports.
- LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics.
- On an in-house benchmark, it scores 76.04, ranking second among four compared systems.
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: hackernews | Overall 6.1/10 | Corroboration: 1
Signal 8.4
Novelty 6.2
Impact 2.6
Confidence 7.5
Actionability 3.5
Summary: Hi HN,
After building for 7 months, I'm excited to announce Paperweight V1.
The idea is that your inbox knows who you ever interacted with and knows where your data lives.
- What happened: Hi HN,
After building for 7 months, I'm excited to announce Paperweight V1.
The idea is that your inbox knows who you ever interacted with and knows where your.
- Why it matters: Hi HN,
After building for 7 months, I'm excited to announce Paperweight V1.
The idea is that your inbox knows who you ever interacted with and knows where your.
- What to do: Track for corroboration and benchmark data before adopting.
Deep
Context
Hi HN,
After building for 7 months, I'm excited to announce Paperweight V1.
The idea is that your inbox knows who you ever interacted with and knows where your data lives.
What's new
It never felt right that with a lot of alternatives, in order to reclaim your privacy you first had to hand over your data.
Key details
- Every account you ever created, every service you signed up for, and every online purchase is connected to your email.
- Most people have hundreds of accounts they've forgotten about, creating security risks and privacy exposure.
Paperweight scans your email history locally to find old accounts, companies, mailing lists, data breaches, and personal information.
- You can then unsubscribe, clean up unwanted email, send privacy/GDPR requests, and keep track of your digital footprint.
The main constraint from the start was privacy.
- It never felt right that with a lot of alternatives, in order to reclaim your privacy you first had to hand over your data.
Results & evidence
- Hi HN,
After building for 7 months, I'm excited to announce Paperweight V1.
The idea is that your inbox knows who you ever interacted with and knows where your data lives.
- I haven't been a customer with them for 8+ years, but they still had all my data in their systems.
Happy to answer any questions or hear what you think.
GitHub: http...
Limitations / unknowns
- Most people have hundreds of accounts they've forgotten about, creating security risks and privacy exposure.
Paperweight scans your email history locally to find old accounts, companies, mailing lists, data breaches, and personal information.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: hackernews | Overall 5.7/10 | Corroboration: 1
Signal 8.4
Novelty 4.0
Impact 2.6
Confidence 7.5
Actionability 3.5
Summary: Built this as a hands-on lab for teaching web application security.
- What happened: Built this as a hands-on lab for teaching web application security.
- Why it matters: Built this as a hands-on lab for teaching web application security.
- What to do: Track for corroboration and benchmark data before adopting.
Deep
Context
Built this as a hands-on lab for teaching web application security.
What's new
Built this as a hands-on lab for teaching web application security.
Key details
- It is a fictional electricity utility's customer portal with 48 deliberately planted flags (most of OWASP Top 10) you find and exploit CTF-style, including a simulated (rule-based, not a real model) AI assistant with its own prompt-injection-style vuln...
Results & evidence
- It is a fictional electricity utility's customer portal with 48 deliberately planted flags (most of OWASP Top 10) you find and exploit CTF-style, including a simulated (rule-based, not a real model) AI assistant with its own prompt-injection-style vuln...
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: rss | Overall 4.0/10 | Corroboration: 1
Signal 7.3
Novelty 4.0
Impact 2.0
Confidence 3.0
Actionability 5.2
Summary: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
- What happened: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
- Why it matters: Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
- What to do: Track for corroboration and benchmark data before adopting.
Deep
Context
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
What's new
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
Key details
- Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
Results & evidence
- Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.