OpenAI Private Safety Processing Detects Misuse Without Retaining Data

OpenAI is testing Private Safety Processing with eligible enterprise and API customers. Automated systems correlate activity across related interactions and return narrowly defined misuse signals instead of exposing prompts or responses, preserving the company's Zero Data Retention commitments.

#AISecurity #LLMSecurity #OpenAI #SafetyMonitoring #MonitoringAndOperations

https://www.csoonline.com/article/4212398/openai-adds-an-ai-safety-layer-to-detect-misuse-without-retaining-enterprise-data.html
👍1
AWS Enforces Agent Access Controls Outside the Agent

AWS shows how Amazon Bedrock AgentCore agents propagate user authorization context so downstream services enforce access controls. The agent acts as an orchestrator, not a gatekeeper, and restrictions hold even if the agent is manipulated through prompt injection.

#AISecurity #AgenticAI #AWS #Bedrock #IdentityAndAccessManagement

https://www.helpnetsecurity.com/2026/08/20/aws-ai-agents-access-controls/
isolated-vm Type Confusion Enables V8 Sandbox Escape to Host RCE

Endor Labs found a type confusion in the C++ glue code of isolated-vm, a JavaScript sandbox library with over one million weekly downloads used by n8n, Mastra, and other AI agent frameworks. A guest can abuse the transferList to escape the V8 sandbox and hijack the host's control flow.

#AISecurity #SandboxEscape #NodeJS #LLMTools #VulnerabilityManagement

https://www.csoonline.com/article/4212151/critical-flaw-patched-in-popular-javascript-sandbox-used-in-ai-projects.html
procoder v1.3.0 Lets Repositories Own Agent Commit Gates

procoder v1.3.0 lets each repository define its own tools, thresholds, and templates for AI coding agents. A repository names a tool rather than a binary and argv, and the print-don't-write contract stays a guarantee because every candidate tool is tested for formatted output on stdout.

#AISecurity #CodingAgents #GitHub #CommitGate #ApplicationSecurity

https://github.com/azrtydxb/procoder/releases/tag/v1.3.0
Claude Opus 5 Sidesteps Obfuscation Instead of Defeating It

Quarkslab handed sandboxed coding agents progressively hardened AArch64 binaries and asked them to recover hidden strings. Claude Opus 5 never fully deobfuscated a protection; it lifted code snippets into Python, executed them under emulation, and once treated an answer-key file as ground truth.

#AISecurity #ReverseEngineering #LLM #Obfuscation #ApplicationSecurity

https://cybersecuritynews.com/claude-opus-5-routes/
COPA: Continual Preference Optimization for Prompt Injection Defense

COPA treats prompt-injection defense as a lifelong learning problem. Instead of static alignment objectives or attack-specific filters, continual preference optimization adapts the model to evolving attack distributions without sacrificing robustness to previously encountered threats.

#AISecurity #PromptInjection #LLMDefense #Research #ApplicationSecurity

https://arxiv.org/abs/2608.19982
👍2
EchoCoT Extracts Hidden Chain-of-Thought From Black-Box Models

EchoCoT exploits a reasoning replay surface between tool calls to extract hidden chain-of-thought traces near-verbatim from black-box reasoning models. On open-source models it reaches up to 66.4 percent near-verbatim extraction success using API-returned fidelity signals.

#AISecurity #ChainOfThought #ModelExtraction #LLM #AISecurityGovernanceAndAssurance

https://arxiv.org/abs/2608.20055
TrustRAG Certifies RAG Documents via Zero-Knowledge Committee Scoring

TrustRAG adds a committee of domain experts that certifies documents through a zero-knowledge protocol before retrieval. Hidden scores are combined via secure multi-party computation, so tampered or manipulated documents cannot reach LLM outputs in healthcare, finance, or legal settings.

#AISecurity #RAG #ZeroKnowledge #LLMIntegrity #AISecurityGovernanceAndAssurance

https://arxiv.org/abs/2608.20097
👍2
Join our 9300+ followers on X

https://x.com/aisechub
NIST Cybersecurity Framework 2.0: Quick-Start Guide for Using Artificial Intelligence (AI) for CSF Analysis and Reporting

https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1353.ipd.pdf
👍2
Maybe the “It’s AI, not us” claims will stop here. Maybe
🔥1