Source: arxiv | Overall 6.6/10 | Corroboration: 1
Signal 9.4
Novelty 5.1
Impact 2.0
Confidence 9.5
Actionability 6.5
Summary: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level.
- What happened: We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
- Why it matters: arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...
What's new
arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.
Key details
- However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
- This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
- We introduce SHOVIR, a benchmark for evaluating vision shortcut behavior in RRG.
- SHOVIR extends two spatially annotated chest X-ray datasets, MIMIC-CXR and PadChest-GR, with per-box CheXpert labels, and defines image-level and disease-level occlusion experiments that contrast baseline performance on clean images against localized, regio...
Results & evidence
- arXiv:2606.30201v2 Announce Type: replace-cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness.
- Computer Science > Computer Vision and Pattern Recognition [Submitted on 29 Jun 2026 (v1), last revised 22 Jul 2026 (this version, v2)] Title:SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation View PDF HTML (experimen...
- Submission history From: Filippo Ruffini [view email][v1] Mon, 29 Jun 2026 12:17:35 UTC (1,063 KB) [v2] Wed, 22 Jul 2026 13:09:26 UTC (1,063 KB) References & Citations Loading...
Limitations / unknowns
- However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
- This allows models to achieve competitive scores by exploiting learned priors or spurious correlations, a failure mode we refer to as vision shortcut.
- Comparing predictions across these conditions isolates two failure modes at the disease-class level: direct shortcuts, where a finding persists after its visual evidence is removed, and contextual shortcuts, where detection degrades once co-occurring pathol...
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: arxiv | Overall 6.4/10 | Corroboration: 1
Signal 9.4
Novelty 5.1
Impact 2.0
Confidence 8.7
Actionability 6.5
Summary: arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex.
- What happened: This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation.
- Why it matters: arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
Submission history From: Bingru Li [view email][v1] Thu, 5 Feb 2026 09:55:19 UTC (1,414 KB) [v2] Tue, 21 Jul 2026 17:59:46 UTC (1,463 KB) Current browse context: cs.CL References & Citations Loading...
What's new
arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.
Key details
- While Large Language Models (LLMs) show promise, a significant gap remains between the theoretical capability of LLMs and their practical utility for researchers.
- This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation.
- The platform comprises an Annotator and an optional Reviewer to simulate a peer-review process.
- This platform supports comparative experiments across three main paradigms: Prompt Engineering (Zero-shot/Few-shot/Chain-of-thought), Retrieval-Augmented Generation, and Fine-tuning.
Results & evidence
- arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification.
- Computer Science > Computation and Language [Submitted on 5 Feb 2026 (v1), last revised 21 Jul 2026 (this version, v2)] Title:LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation View PDF HTML (experimental)A...
- Submission history From: Bingru Li [view email][v1] Thu, 5 Feb 2026 09:55:19 UTC (1,414 KB) [v2] Tue, 21 Jul 2026 17:59:46 UTC (1,463 KB) Current browse context: cs.CL References & Citations Loading...
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.
Source: arxiv | Overall 6.2/10 | Corroboration: 1
Signal 9.4
Novelty 4.0
Impact 2.0
Confidence 8.7
Actionability 6.5
Summary: arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and.
- What happened: We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation.
- Why it matters: arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation.
- What to do: Validate with one small internal benchmark and compare against your current baseline this week.
Deep
Context
Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning.
What's new
arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user i...
Key details
- It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video.
- Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response.
- We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p.
- Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts.
Results & evidence
- arXiv:2607.18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user i...
- We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p.
- We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk.
Limitations / unknowns
- Generalization outside curated tasks is still unclear.
Next-step validation checks
- Reproduce one claim with a public baseline and fixed evaluation settings.
- Check robustness on out-of-distribution or long-context cases.
- Track whether independent teams report matching results.