Source: arxiv | Overall 6.5/10 | Corroboration: 1
Signal 9.4
Novelty 5.1
Impact 2.0
Confidence 9.5
Actionability 6.5
Summary: arXiv:2609.08355v1 Announce Type: cross Abstract: Solving repository-level code tasks requires LLM-based agents to use code search tools to navigate large codebases and identify a.
- What happened: We introduce RepoNav, a lightweight post-retrieval interface that reorganizes retrieved snippets into a file-centered navigation scaffold.
- Why it matters: Across diverse models on LocBench, RepoNav improves function-level localization and narrows the file-to-function gap.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
arXiv:2609.08355v1 Announce Type: cross Abstract: Solving repository-level code tasks requires LLM-based agents to use code search tools to navigate large codebases and identify a small set of relevant files and functions.
What's new
Controlled ablations demonstrate that these gains come from structured evidence organization rather than simply exposing additional file structure, and the approach also improves performance on a repository-level question-answering benchmark.
Key details
- However, current retrieval tools typically return flat lists of isolated code snippets: such lists can surface relevant files, but provide insufficient structure for agents to distinguish the target function from semantically similar alternatives in the sam...
- We introduce RepoNav, a lightweight post-retrieval interface that reorganizes retrieved snippets into a file-centered navigation scaffold.
- By presenting compact structural cues and candidate targets, this scaffold guides on-demand file-structure browsing, helping agents compare sibling symbols before selecting a target function.
- Across diverse models on LocBench, RepoNav improves function-level localization and narrows the file-to-function gap.
Results & evidence
- arXiv:2609.08355v1 Announce Type: cross Abstract: Solving repository-level code tasks requires LLM-based agents to use code search tools to navigate large codebases and identify a small set of relevant files and functions.
- Computer Science > Software Engineering [Submitted on 8 Sep 2026] Title:RepoNav: From Snippet Retrieval to File-Centered Repository Navigation for Code Agents View PDF HTML (experimental) Abstract:Solving repository-level code tasks requires LLM-based agent...
Limitations / unknowns
- However, current retrieval tools typically return flat lists of isolated code snippets: such lists can surface relevant files, but provide insufficient structure for agents to distinguish the target function from semantically similar alternatives in the sam...
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: arxiv | Overall 6.3/10 | Corroboration: 1
Signal 9.4
Novelty 4.0
Impact 2.0
Confidence 9.5
Actionability 6.5
Summary: arXiv:2609.08790v1 Announce Type: cross Abstract: Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into.
- What happened: To address these gaps, we introduce AHLERT, a system that automatically extracts relevant, environment-aware, and hunt leads from threat reports through (i) a hybrid.
- Why it matters: arXiv:2609.08790v1 Announce Type: cross Abstract: Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
arXiv:2609.08790v1 Announce Type: cross Abstract: Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into actionable hunt leads: concise, investigable hypotheses grounded in observable artifact...
What's new
Existing automated approaches stop at the entity layer, ignore the defender's operational environment, and analyze each report in isolation.
Key details
- Producing such leads manually is a tedious and hard-to-scale task.
- Existing automated approaches stop at the entity layer, ignore the defender's operational environment, and analyze each report in isolation.
- To address these gaps, we introduce AHLERT, a system that automatically extracts relevant, environment-aware, and hunt leads from threat reports through (i) a hybrid retriever that combines dense vector search with multi-hop traversal over a knowledge graph...
- We evaluate AHLERT on public CTI reports for well-known APTs across multiple proprietary and open-weight models.
Results & evidence
- arXiv:2609.08790v1 Announce Type: cross Abstract: Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into actionable hunt leads: concise, investigable hypotheses grounded in observable artifact...
- Hybrid evidence retrieval with ontology grounding raises mean F1 by ~2x (0.44 to 0.85) over a single-route flat-RAG baseline, and AHLERT attains the highest effectiveness score (~86.95%) compared with off-the-shelf LLM models.
- Computer Science > Cryptography and Security [Submitted on 8 Sep 2026] Title:Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports View PDF HTML (experimental) Abstract:Threat hunting increasingly depends on converting unstruct...
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: arxiv | Overall 6.4/10 | Corroboration: 1
Signal 9.4
Novelty 5.1
Impact 2.0
Confidence 8.7
Actionability 6.5
Summary: arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by.
- What happened: arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many.
- Why it matters: On the six fresh problems of the 2026 International Mathematical Olympiad, every model with room to improve scores higher under AutoFyn than in its provider's own coding.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
On the six fresh problems of the 2026 International Mathematical Olympiad, every model with room to improve scores higher under AutoFyn than in its provider's own coding agent.
What's new
arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights.
Key details
- Each round begins from a fresh model session, and durable information is reintroduced only through explicit interfaces such as persistent memory files, reports, and repository state.
- Within a round, an orchestrator explores, plans and builds many alternative approaches with specialized agents, while a task-grounded verifier verifies the work and supplies an objective reward for measuring progress.
- This reward is distilled back into the persistent state, which updates the effective policy for the next round.
- In this technical report, we formalize this loop and describe its persistent state and verification interfaces.
Results & evidence
- arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights.
- On the six fresh problems of the 2026 International Mathematical Olympiad, every model with room to improve scores higher under AutoFyn than in its provider's own coding agent.
- AutoFyn also built the top-ranked agent on the Spider 2.0 dbt benchmark, and has produced $16$ maintainer-confirmed vulnerability advisories in Next.js, MetaMask, pnpm, Warp, LiteLLM, Langflow, and Open WebUI.
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.