Loading...
Loading...
Browse, search, and filter preprints from arXiv—fast, readable, and built for curious security folks.
Showing 18 loaded of 52,533—scroll for more
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors assess subtasks in isolation. To address this gap, we introduce PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment. For training and evaluation, we construct a Multi-Step Plan Safety (MSP-Safe) dataset through paired task construction, plan generation using diverse planners, and safety annotation by three judges. Task-oriented SFT on MSP-Safe establishes fundamental plan-safety assessment capabilities, yet a substantial gap remains between compact models suitable for real-time deployment and stronger but costlier large models. Accordingly, we propose Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision along their on-policy trajectories. It combines token-level distribution transfer from a fine-tuned strong teacher with probability-routed sequence-level compensation, retaining student-generated targets when the student favors the reference safety decision and using teacher-reconstructed targets otherwise. Across all test subsets, PlanGuard-2B achieves average 87.15% ACC and 87.21% F1, demonstrating effective whole-plan physical-risk detection at compact model scale. Code and dataset will be publicly released.
As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way that is aligned with the users intent. Auditing an agent's network traffic provides a clear record of the agent interactions. This work presents a novel approach to monitoring the network traffic of agentic systems using complex valued hypersparse traffic matrices by integrating DBOS (DataBase OS), the OneSparse PostgreSQL database, and the GraphBLAS math library. To develop these concepts an agentic simulator was constructed, allowing a varying numbers of AI agents to collectively survey a virtual environment using different strategies. The resulting network traffic matrices enable easy monitoring of the AI agents.
Deepfakes have rapidly emerged as a pressing threat to information integrity and security because they exploit human trust in visual and auditory perception. Yet, little is known about whether humans and their underlying (sub)conscious neuro-physiological processes can reliably distinguish deepfake from real videos. We introduce DECEIVE (Deepfake Exploitation of Cognitive Engagement and Implicit Visual Evaluation), a framework that models how deepfake videos are validated as adversarial payloads through behavioral and neuro-physiological screening of viewers, and how attacks can be refined by selecting payloads that evade detection. The framework is dataset agnostic and applies to synthetic or real media. It is inherently dual-use: an adversary with equivalent measurements could iterate on candidate manipulations and retain those that evade human detection. This motivates open, defensive evaluation. Measuring which deepfakes defeat human perception establishes a realistic bound on attacker capability against which detection tooling, provenance and watermarking mechanisms, and user-facing protections can be assessed. As an instantiation, we conducted an EEG and eye-tracking study in which participants viewed real, deepfake, and look-alike videos drawn from Celeb-DF and a curated celebrity set, while behavioral judgments and implicit responses were recorded. Contrary to expectations of subconscious differentiation suggested by prior work on paintings and phishing websites, no statistically significant neuro-physiological differences emerged between real and deepfake videos, although clear distinctions were observed for look-alike videos. Behaviorally, participants accepted 26.68% of manipulated clips as authentic, rising to 31.94% for familiar identities, confirming the studied deepfakes as effective adversarial payloads within DECEIVE.
Detecting distributed denial-of-service (DDoS) attacks in cloud-integrated IoT networks is difficult when labeled traffic is scarce. Generative semi-supervised learning can supplement the available training data, but prediction shifts induced by synthetic views may affect the targets assigned to real unlabeled flows. We propose AnchorMixGAN, a generative semi-supervised framework that addresses this problem through anchor-aligned target construction. Its Anchor-MAS module treats each real unlabeled flow as an anchor and creates alternative views by replacing one field group at a time with values from generated traffic. A frozen reference classifier predicts the anchor and its views; averaging and sharpening these predictions produces a soft target for the original flow. The flow and its target are then mixed with a labeled example using MixUp, allowing the detector to learn from both the original labeled records and the mixed examples. We analyze how reference-classifier error, view construction, and sharpening affect the target, and derive a bound on the resulting change in cross-entropy at a fixed detector prediction. At the reported 90% training setting with 20% of the training records labeled, AnchorMixGAN attains accuracies of 97.3%, 97.4%, and 96.5% on NSLKDD, BoT-IoT, and CICIoT2023, respectively, exceeding the corresponding MixGAN results by 1.6, 1.0, and 4.4 percentage points.
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
Cloud-based large language models (LLMs) require users to disclose plaintext data to service providers, creating privacy risks in sensitive domains. Existing privacy-preserving approaches often trade utility for protection, incur substantial computational or communication overhead, remain vulnerable to reconstruction from intermediate representations, or protect only a subset of the training and inference pipeline. We introduce Client-Resolved Generation (CRG), a genera- tion interface that separates server-side generation from the lexical realization of input-derived content. The client transmits only pooled and noise-perturbed rep- resentations, while input-derived output content is represented using request-local positional references and resolved to its original strings only on the client. This interface protects private input and input-derived output content during both train- ing and inference while allowing the service provider to keep its proprietary model parameters hidden from the client. At the same time, exact lexical reuse remains possible without directly exposing the reused content on the provider-visible gen- eration path. We evaluate CRG on medical and document-grounded QA, sensi- tive identifier transfer, and tool calling, together with reconstruction and raw-logit leakage analyses. On SealTools, CRG improves complete-call exact match from 57.3% to 79.9% over the input-privacy framework PPFT, with larger gains as more required output content can be resolved through references. Together, these results show that CRG provides a practical interface for privacy-sensitive cloud LLMs by reducing plaintext exposure across both input and output pathways while preserv- ing task utility and server-side model confidentiality.
LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that evaluation is rarely examined. We audit an IPI benchmark and its harness and identify four defect classes -- silent payload non-delivery, attack success scored by tool identity rather than arguments, false-rejection rate conflated with model incapacity, and the absence of an audit trail -- each of which yields a plausible, publishable, and incorrect number. We quantify the distortion by re-scoring identical execution traces under the defective and corrected definitions: on real agent behaviour, the tool-identity scorer reports a 21.7% attack-success rate where the true argument-level rate is 1.2%. In the sharpest case, an open model previously reported at 62.8% registers 0% under the corrected harness -- the prior figure largely an artifact of undelivered payloads and identity-level scoring. We release a harness whose construction makes each defect unrepresentable -- machine-checkable payload placement, argument-level attacker predicates, per-scenario environments, and mandatory trace persistence -- and use it to report three quantities the field does not: whether a compromised agent discloses the attack, the full security/utility operating curve of an LLM-judge defense, and tool-calling capability disentangled from defensive over-blocking. A corrected harness further overturns a reported "capability barrier": a model deemed incapable of tool use is in fact fully capable, its earlier result an artifact of environment mismatch. We argue that evaluation validity is a prerequisite for, not a footnote to, defense claims in agentic security, and provide an instrument that enforces it.
As local small language models (SLMs) increasingly collaborate with more capable cloud large language models (LLMs), a natural privacy question arises: Can a local SLM obtain cloud LLM guidance while protecting user privacy? Existing privacy-preserving SLM-LLM frameworks primarily hide sensitive values while preserving task semantics, which can still expose what the user is trying to accomplish. For example, allocating scarce medical supplies across hospitals may signal an emerging public-health emergency, while rebalancing an investment portfolio may reveal a private investment strategy, even when names and numerical values are hidden. Recent decoy-based methods further obscure task intent by hiding the real request among alternatives, but stronger protection relies on more decoys or semantic abstraction, increasing overhead or risking utility loss. More fundamentally, existing work does not systematically characterize the components of private task intent or how each should be protected. We therefore introduce task-private consultation, which characterizes task intent through two components: task context and task operation. To the best of our knowledge, this is the first systematic study of these components and their individual and joint protection in local-cloud SLM-LLM consultation. To realize this setting, we propose PriCon, an end-to-end framework that transforms the task itself through recoverable mathematical reformulation rather than hiding it among alternatives. A local closed-loop refinement mechanism further maintains privacy and recoverability throughout consultation. Experiments on 100 tasks show that PriCon reduces cloud-side task-intent inference Hit@1 to nearly 0%, versus 93-99% under sensitive-value removal and 3-30% under decoy-based protection, while preserving cloud-assisted utility.
LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents display, and an attacker can spoof them. Displayed identity thus decides operational authority, meaning who is trusted to check the work and who is allowed to change it. As a result, a risky subagent can keep authority over execution even after other evidence contradicts it. We introduce TrustFork, an LLM agent safety benchmark with 1,890 tasks and 27,826 valid trajectories across 16 agent systems. These systems run eight orchestrators under the OpenCode, OpenClaw, and Pi harnesses. In each task, one subagent carries a risky goal while the other three stay aligned with the user, so contradicting evidence can exist. A task can also change the identity a subagent displays without changing the model behind it, which lets us trace a shift in authority to the label. Our analysis shows that even when another subagent contradicts the risky response, the orchestrator still acts on it in 72.0% of cases on average, most often in the systems with the least terminal harm. Swapping the family labels nearly triples how often the orchestrator obtains the risky response. A safer response is available in 84.0% of tasks, yet it decides the outcome in only 25.0%. The harness also decides which responses reach the orchestrator. Among three runtime defenses, hiding identity cues helps most consistently, while verifying before action helps only when the harness returns enough evidence. TrustFork shows that production agent orchestration must bind authority to evidence before execution causes harm. Our project is in https://henrymao2004.github.io/agent-orchestration-safety/.
LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, and destructive over-correction when acting on it damages previously correct work. We introduce CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks. Every scored run first reaches a verified correct state, whose supporting rationale and history stay in the workspace while the facts that would settle the accusation lie in external or runtime state beyond the agent's reach. The agent cannot confirm or refute the claim with a local check, so the right response should keep the work and ask for the missing evidence. Each task either hands the agent correct work with saved evidence or let it build and verify that work first, and five risk factors set how the accusation enters the workflow. We score accusation acceptance and evidence use from the trajectory and measure harm by deterministic replay of downstream events. Across 14 of the latest models in Claude Code, false accusations damage correct work in up to 60.06% of runs, and stronger models often do so after recovering the supporting evidence. The same model behaves differently across OpenCode, Codex, and Hermes, and a harness gate driven by the benchmark's live signals cuts replayed harm by 74%. These results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents. Our project is in https://henrymao2004.github.io/agent-over-correction/.
Networks of interacting Hebbian networks have recently been shown to perform a task beyond associative memory, namely \emph{pattern disentanglement}: when fed with a spurious mixture of stored patterns, the different modules spontaneously specialize on, and retrieve, the different constituents of the mixture. So far, this capability has only been established for pairwise interactions, which limits the number of patterns that can be handled. Here we introduce a dense extension of these modular networks, in which both the auto-associative couplings within each module and the hetero-associative couplings among modules are promoted to higher-order Hebbian interactions. We show that, with a suitable choice of the interaction orders, the network disentangles mixtures while storing a number of patterns that scales linearly with the module size, a regime where its pairwise counterpart fails. Through a statistical-mechanical analysis based on Guerra's interpolation, we derive the self-consistency equations for the order parameters and draw the phase diagrams identifying the region where disentanglement is achieved; these predictions are confirmed by Monte Carlo simulations. Finally, we show that disentanglement provides a natural decoding primitive, and we illustrate it with two applications: the explicit reconstruction of all the hidden patterns from the Hebbian tensors and a stream of unlabeled mixtures, and a proof-of-concept communication protocol in which each message token is transmitted as a masked mixture of hidden patterns and decoded by the network dynamics. Owing to its attractor-based decoding, the protocol degrades gracefully under strong channel corruption, where conventional secure-transmission pipelines fail abruptly.
Quantum information enables many cryptographic primitives that are impossible in the classical world. A line of works has developed cryptographic protocol under the assumption that quantum adversaries are restricted to noisy intermediate-scale quantum (NISQ) computing power, enabling strong one-time functionalities. But the advent of early fault-tolerant quantum computers eras will allow deeper logical quantum circuits, calling into questions the applicability of these NISQ-based assumptions. In this work, we adapt the classically accessible random oracle model (CAROM) as in [BDF+11] and [AK22], in which adversaries are only allowed to classically query the random oracle. The restriction is well motivated for NISQ quantum adversaries and may remain plausible in the presence of early fault-tolerant quantum computers. Then, we show that an efficient simulation-secure one-time memory (OTM) is possible under CAROM. Our protocol uses only BB84 states and has quadratic communication: for a $λ$-bit message and integer-valued parameters $n=n(λ)$ and $\ell=\ell(λ)$, the construction uses $n\ell$ qubits and $(n+2)λ$ classical bits, and for any quantum adversary with at most $2^{\ell/2-1}-1$ classical queries to the random oracle, its simulation advantage is at most $(n+3)\left(\frac{3}{4}\right)^n.$ Hence exponentially small simulation advantage in $n$.
Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability. Instead of scoring the verdict, we score whether the agent retrieved the code its conclusion depends on. We present VulContextBench, a benchmark of 111 vulnerability-introducing commits (VICs) across 83 repositories, 63 CWEs, and five languages. Existing datasets label such commits by tracing a fix back through the version history, which often points to the wrong commit. We therefore audit every case by hand against an explicit four-criterion definition of a VIC, so the benchmark does not inherit that label noise. Each case is annotated with gold context, 464 code blocks in total, each tagged by its role in the evidence for the vulnerability. We evaluate seven frontier models with precision, recall and F1 at three granularities (file, block, and line), scored separately on the context an agent viewed while exploring and on the context it finally declared as evidence. The gap between the two is the main finding. Every model opens most of the gold context while exploring, but reports only part of it as evidence. At the level of code blocks, the share of the gold context a model reports is 37 to 73 percentage points below the share it viewed. Qwen3-Coder-Next views 86.3% of the lines in annotated code blocks but cites only 12.9% in its final report. GPT-5.5, which cites the most, views 73% and reports 36%. These results highlight a gap between finding relevant code and selecting it for the final report, which verdict-level benchmarks cannot reveal.
Anyone with a public footprint leaks facts that were never stated, and language models make the inference cheap. We present a framework for measuring and reducing this inference exposure that runs on the owner's own CPU with no language model at analysis time, instantiated on organisations and on individuals. It starts from a measurement result: scoring an inference system against the target's private truth conflates how well the system reads the record with how much the record leaks. On a 128-question instrument over sixteen synthetic firms, almost half of the questions are never answered correctly by any of six readers, four of them language models, and a majority-class guess accounts for most of every reader's score. We therefore separate reading accuracy from leakage rate and introduce an injection protocol that creates cells with known support. Our analyser combines rules, statistical solvers and a 106M-parameter encoder trained from scratch that marks verbatim evidence and never generates text; every answer carries a graded certificate whose recorded proof replays. Its certified answers are correct in 93% of resolved cases, against 49-73% for the language models' quote-backed answers, whose citations are produced alongside the answer rather than deriving it; with plain-prose articles in the record, 70% of its evidence-bearing answers rest on evidence that establishes them, against 18-56% for the models. A constrained defence that rewrites each fact's carrier as a true but coarser statement hides every single-carrier fact from four language-model adversaries at 40% lower edit cost than deletion. On sixteen synthetic people the guessing term is larger still, and a decoy planner with no language model halves the correct answers of the estimator it targets without transferring to a second.
Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organization-specific, rapidly evolving SOC operational standards. We introduce REFINE, an LLM-agent framework for enterprise alert triage. REFINE encodes analyst expertise as structured skills and continuously adapts using analyst disposition feedback. It enforces recall = 1.0 as a hard constraint during evolution to maximize auto-closure of false positives, and identifies judgment blind spots by combining alert distributions with model error boundaries. Evaluated on four real industrial SOC scenarios across four MITRE ATT&CK phases with temporal split: REFINE achieves recall=1.0 on all evolution sets. On future test windows, it retains recall=1.0 in three scenarios; the degraded case reaches 0.807 recall, still outperforming self-evolution baselines (0.49-0.58).
An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across sixteen deployed frameworks, none writes one in full. Hearsay examines the record, not the task: five harnesses run fourteen tasks, three blinded LLM examiners and a human panel read the records, and every excerpt an examiner quotes is checked mechanically for who wrote it. First, the record lets a reader name the fault but not prove how the run went. Examiners name the right fault in 74 to 91% of 140 runs, but the fault can be proved only from two files the benchmark adds; for what happened in between, fewer than one citation in ten lands on anything the harness did not write, and the examiner with the fewest false alarms catches half of the entries we delete, rewrite or fabricate. Second, the remedy is a second author, not a stronger seal on the first. An append-only log of what passes between harness and model, kept outside the harness and read against the record in both directions, reports all 28 omissions and fabrications we made a harness commit as it ran, where a hash chain over the harness's own record passes all 28. Handed the log, examiners keep their fault verdicts but rest more of their citations on what the harness did not write. What makes a record evidence is who writes it, not what is captured.
Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery, exploitation, and security analysis tasks, leaving the evaluation of attack chain provenance under realistic security logs insufficiently studied. To address this gap, we introduce CyberClear, a benchmark for evaluating LLM agents and advanced agent systems on APT attack chain provenance from long-context security logs. CyberClear covers both single-step attacks and multi-stage attack chains, requiring agents to identify attack evidence, infer attack progression, and generate provenance graphs containing entities, causal relationships, MITRE ATT&CK techniques, and forensic evidence. To enable comprehensive evaluation, we develop an evaluation method tailored to APT attack chain provenance. Unlike conventional text similarity metrics that focus on surface-level matching, our evaluation examines whether reconstructed graphs preserve the semantics of attack chains across single-step behavior correctness, multi-step behavior identification, temporal and causal consistency, entity and relationship fidelity, and overall attack narrative consistency. Advanced multi-agent systems powered by state-of-the-art LLMs still struggle on CyberClear, motivating us to propose CyberProvenance, an agent cyber harness designed for multi-agents that augments LLM agents with evidence accumulation, execution-based validation, and feedback-guided refinement mechanisms for reliable attack-chain provenance. Extensive evaluations on CyberClear demonstrate the effectiveness of CyberProvenance in improving evidence reasoning, execution-grounded validation, and complete APT attack chain reconstruction.
Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill may pass pre-execution scanning yet fail to realize its target under runtime defenses, while a revision that repairs execution may introduce new scanner findings. We introduce SkillDRE, a fully automated framework for evolving complete malicious skill packages through a dual-stage feedback loop. Given a benign task and its associated skills, SkillDRE autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule. It then holds both fixed while evolving the skill implementation, with preservation of legitimate task capability. SkillDRE combines scanner-guided evolution with runtime-guided refinement informed by execution outcomes observed under runtime defense. Each runtime-guided revision returns to the pre-execution stage for rescanning and further optimization before re-execution, forming a cross-stage closed loop. Evaluated on SkillsBench across four victim models, SkillDRE achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance. These results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability. Codes is available at https://github.com/whfeLingYu/SkillDRE