{"data":[{"id":"01M3B85J3C5E4A1Z16ZP41AJZE","url":"https://arxiv.org/abs/2608.02691","title":"Output-Aware Rotation for INT2 KV-Cache Quantization","summary":"The paper proposes a new method for ultra-low-bit quantization of key-value caches in large language model inference, specifically addressing the mismatch between cache optimization and output projection. The method, called OptR, improves performance and preserves the paged KV-cache format with negligible overhead.","source":"rss","tags":["kv-cache-quantization","llm-infrastructure","research","model-performance","optimization"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:07:34.396Z"}}},{"id":"01M3B85HVBHY9XWWMNGKT2TQTF","url":"https://arxiv.org/abs/2607.22115","title":"Benchmarking Text-to-SQL under Role-Based Access Control","summary":"A new benchmarking framework for text-to-SQL systems under role-based access control is proposed, with an LLM-assisted workflow that augments existing benchmarks with realistic user roles and access policies. The framework includes evaluation metrics to identify RBAC-specific failure modes and disentangle SQL utility from access-control compliance.","source":"rss","tags":["llm-security","benchmarking","rbac","access-control","advisory"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:07:34.608Z"}}},{"id":"01M3B85HK01G5YE784GHSGV6ZR","url":"https://arxiv.org/abs/2607.14236","title":"Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection","summary":"A new post-training framework, LIFT, adds contact reactivity to pretrained VLA policies, improving their manipulation capabilities in occluded or ambiguous scenes. LIFT uses a reactive action expert and online DAgger loop to update the policy during execution, leading to faster learning and higher performance.","source":"rss","tags":["vlam","rl","dagger","contact-rich-manipulation"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:07:33.907Z"}}},{"id":"01M3B85FBT073WF5GWWMHYH7AC","url":"https://arxiv.org/abs/2605.16268","title":"Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes","summary":"Researchers developed an LLM-powered triage agent for banking operations to improve customer assistance and case classification. The agent conducts multi-turn conversations, asks questions, and routes cases to specialist teams. The study evaluates the agent's accuracy, robustness, and compliance, achieving a 30.6% increase in classification accuracy and high subject-matter expert satisfaction.","source":"rss","tags":["llm","policy","agent","compliance","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:07:41.703Z"}}},{"id":"01M3B85F2TTG8ZYP227RDPCFXC","url":"https://arxiv.org/abs/2605.11750","title":"DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies","summary":"Researchers propose DreamAvoid, a test-time dreaming framework to help Vision-Language-Action models anticipate and avoid failures in critical phases. The framework uses a Dream Trigger, Action Proposer, and Dream Evaluator to predict and select optimal actions. Evaluations on real-world tasks and simulation benchmarks show a 23.7% increase in success rate compared to the base policy.","source":"rss","tags":["dreamavoid","vla-models","test-time-dreaming","failure-avoidance","vision-language-action","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T04:07:12.500Z"}}},{"id":"01M3B85E6V6H0BXJCC20FWQ3D6","url":"https://arxiv.org/abs/2605.03228","title":"Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory","summary":"Researchers propose ShadowMem, a defensive framework to counter long-horizon threats against LLM agents. This framework maintains a safety-focused memory to assess the risk of pending actions, achieving high detection accuracy and low overhead.","source":"rss","tags":["llm-security","mcp-security","shadow-memory","long-horizon-threats","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"security","classified_at":"2026-09-25T03:07:43.372Z"}}},{"id":"01M3B85DXCP62YV51JD57663G1","url":"https://arxiv.org/abs/2604.13061","title":"Toward Measuring Structural Drift in LLM Communication Loops","summary":"A research paper proposes two metrics for measuring structural drift in LLM communication loops, which occurs when information is dropped, compressed, or misrouted in stateful pipelines. The paper introduces structural communication coherence metrics and demonstrates their effectiveness in revealing directional interaction structures across human-to-human, human-to-LLM, and LLM-to-LLM dialogues.","source":"rss","tags":["llm-security","research","communication-coherence"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:07:50.506Z"}}},{"id":"01M3B85C764M4X3ZNBF9R5Z3A6","url":"https://arxiv.org/abs/2603.18740","title":"Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review","summary":"Researchers studied how Large Language Models (LLMs) in Automated Code Review (ACR) systems can be exploited through contextual bias. They found that attackers can inject bias into ACR judgments, allowing them to reintroduce vulnerabilities into code. This highlights the importance of human oversight and contributor trust in the development process.","source":"rss","tags":["llm-security","supply-chain","research","vulnerability"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"security","classified_at":"2026-09-25T03:07:51.195Z"}}},{"id":"01M3B859NHQVAGM3SYSK6KV8FW","url":"https://arxiv.org/abs/2511.15375","title":"Parameter Importance-Driven Continual Learning for Foundation Models","summary":"A new method, PIECE, is introduced for preserving the general ability of foundation models while learning new domain knowledge without accessing prior training data or increasing model parameters.","source":"rss","tags":["continual-learning","foundation-models","language-models","multimodal-models","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:07:57.240Z"}}},{"id":"01M3B858TD2QAJQ3T2VNFJ28K5","url":"https://arxiv.org/abs/2510.01354","title":"WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents","summary":"A research paper presents the first comprehensive benchmark study on detecting prompt injection attacks targeting web agents. The study introduces a fine-grained categorization of such attacks and evaluates the performance of various detection methods across multiple scenarios, highlighting their strengths and weaknesses.","source":"rss","tags":["llm-security","prompt-injection","web-agents","research","security","benchmarking"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"security","classified_at":"2026-09-25T03:08:07.978Z"}}},{"id":"01M3B8547B3EQ8QPV2ZKN46BZR","url":"https://arxiv.org/abs/2607.11433","title":"Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents","summary":"A new evidence-ledger planning approach, Omni-Decision, is presented for omni-modal agents. It records and manages evidence to improve decision-making and reduce the impact of noisy observations. This innovation is relevant to developers of AI agents, as it addresses a key challenge in multimodal agent planning.","source":"rss","tags":["omni-modal-agents","evidence-ledger","planning","research","agent-relevance"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:08:17.256Z"}}},{"id":"01M3B8530ZXRWS6145W5WKH1Y1","url":"https://arxiv.org/abs/2605.26114","title":"MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research","summary":"MobileGym is a browser-hosted simulation platform for mobile GUI agent research, enabling verifiable outcome signals and scalable online RL through parallel rollouts. It captures and configures environment state as structured JSON, allowing for practical state programmability and task creation at scale.","source":"rss","tags":["research","parallel-rl","gui-agent-research","simulation-platform","rl"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:08:23.176Z"}}},{"id":"01M3B850J6EE06PF52760JMAE2","url":"https://arxiv.org/abs/2609.28449","title":"Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark","summary":"A repository-level dynamic benchmark for evaluating large language model (LLM) ability to reason about code execution is introduced. The benchmark, called SWE-Flux, contains 480 execution-grounded instances across 12 real Python repositories, and evaluates LLMs on control flow, loops, program state, dataflow, exceptions, and program invariants. Results show that current LLMs struggle with precise state reasoning and inter-procedural execution.","source":"rss","tags":["llm-infrastructure","ai-research","model-evaluation","research","benchmark"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:08:39.304Z"}}},{"id":"01M3B8500JENJ9SDVYCB5QNTS0","url":"https://arxiv.org/abs/2609.28416","title":"Agent-Editing World Model: Rethinking World Modeling for LLM Agents","summary":"The Agent-Editing World Model (AEWM) is proposed for improving the performance of large language model (LLM) agents. AEWM models how reasoning and actions shape future task progress, addressing task-state contamination by allowing agents to edit their own reasoning and actions. AEWM combines Action Judge and State Revision capabilities with real execution, and is trained across various domains with improved results compared to existing models.","source":"rss","tags":["research","llm-security","agent-security","ai-infrastructure","model-architecture"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:08:38.965Z"}}},{"id":"01M3B84YX21Y332R6DPHHGFTGW","url":"https://arxiv.org/abs/2609.28372","title":"Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer","summary":"Researchers analyzed how Large Language Models (LLMs) used by shopping agents make purchasing decisions. They found that when LLMs are given a vague goal prompt, they omit diagnostic attributes and make suboptimal choices, resembling human heuristics. This highlights the importance of storefront information architecture in governing AI shopping behavior.","source":"rss","tags":["llm-security","mcp-security","agent-security","advisory"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:08:34.969Z"}}},{"id":"01M3B84Y2G1CMZXNBWQ8BDBWQN","url":"https://arxiv.org/abs/2609.28256","title":"MemBodied: Recurrent Associative Memory for Vision-Language-Action Models","summary":"MemBodied is a fixed-size episodic memory for Vision-Language-Action models that retains past observations in context to improve performance in history-dependent manipulation tasks. It achieves higher success rates than stateless policies and memory-augmented baselines, with fewer added parameters.","source":"rss","tags":["llm-research","agent-relevant","memory-augmentation","vision-language-action-models","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:08:46.551Z"}}},{"id":"01M3B84WZTVDA1YY2XXQEVTNZN","url":"https://arxiv.org/abs/2609.28216","title":"From Agent Output to Authorized Transition","summary":"This paper proposes the Agile-V Assurance Spine, a cross-domain transition contract for software, firmware, and PCB engineering. It presents a framework for evidence-gated lifecycle control, continuous assurance, runtime admission, provenance, and AI/ML inventories.","source":"rss","tags":["agile-v-assurance-spine","transition-contract","evidence-gated-lifecycle-control","continuous-assurance","ai-ml-inventories","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:08:47.921Z"}}},{"id":"01M3B84RCH6SGSXNDXDABPP995","url":"https://arxiv.org/abs/2609.27891","title":"Schr\\\"odinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?","summary":"Researchers propose Schr\"odingerRepo, an evaluation framework for testing coding agents by dynamically instantiating repository representations, which helps to prevent memorization of canonical repository cues and assesses robust repository reasoning.","source":"rss","tags":["llm-security","research","evaluation-framework","agent-testing","repository-reasoning"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"high","category":"technology","classified_at":"2026-09-25T03:09:02.429Z"}}},{"id":"01M3B84PCYPAN6W4B2Z7HNW24V","url":"https://arxiv.org/abs/2609.27845","title":"Query Implied Generative Engine Optimization","summary":"A new approach to Generative Engine Optimization (GEO) called Query Implied Generative Engine Optimization (QI-GEO) infers user intent from documents to improve content visibility in generative search settings, using Large Language Models (LLMs).","source":"rss","tags":["llm","generative-search","geo","query-implied-geo","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"medium","category":"technology","classified_at":"2026-09-25T03:09:11.759Z"}}},{"id":"01M3B84P4N8E6YRN7KEA2HH03R","url":"https://arxiv.org/abs/2609.27828","title":"A Non-Invasive Cloud-Based Migration Strategy for Post-Quantum Cybersecurity in Smart HVAC Systems: Architecture, Implementation, and Empirical Evaluation","summary":"A research paper proposes a non-invasive post-quantum cybersecurity solution for smart HVAC systems, using a PQC proxy on a Raspberry Pi gateway to perform ML-KEM-768 encapsulation and ML-DSA-65 authentication.","source":"rss","tags":["pqc-proxy","rv","implementation","research"],"created_at":"2026-09-24T04:00:00.000Z","ingested_at":"2026-09-25T03:00:19.175Z","score":0,"status":"live","metadata":{"source_id":"01KZH5DJ583SFWTH571YFWTKN8","classification":{"model":"@cf/meta/llama-3.1-8b-instruct-fp8","path":"fallback","relevance":"low","category":"technology","classified_at":"2026-09-25T03:09:11.002Z"}}}],"meta":{"total":650,"limit":20,"offset":0}}