OpenAI Private Safety Processing Detects Misuse Without Retaining Data
OpenAI is testing Private Safety Processing with eligible enterprise and API customers. Automated systems correlate activity across related interactions and return narrowly defined misuse signals instead of exposing prompts or responses, preserving the company's Zero Data Retention commitments.
#AISecurity #LLMSecurity #OpenAI #SafetyMonitoring #MonitoringAndOperations
https://www.csoonline.com/article/4212398/openai-adds-an-ai-safety-layer-to-detect-misuse-without-retaining-enterprise-data.html
OpenAI is testing Private Safety Processing with eligible enterprise and API customers. Automated systems correlate activity across related interactions and return narrowly defined misuse signals instead of exposing prompts or responses, preserving the company's Zero Data Retention commitments.
#AISecurity #LLMSecurity #OpenAI #SafetyMonitoring #MonitoringAndOperations
https://www.csoonline.com/article/4212398/openai-adds-an-ai-safety-layer-to-detect-misuse-without-retaining-enterprise-data.html
CSO Online
OpenAI adds an AI safety layer to detect misuse without retaining enterprise data
The new system is designed to detect AI misuse across interactions without exposing prompts or responses.
👍1
AWS Enforces Agent Access Controls Outside the Agent
AWS shows how Amazon Bedrock AgentCore agents propagate user authorization context so downstream services enforce access controls. The agent acts as an orchestrator, not a gatekeeper, and restrictions hold even if the agent is manipulated through prompt injection.
#AISecurity #AgenticAI #AWS #Bedrock #IdentityAndAccessManagement
https://www.helpnetsecurity.com/2026/08/20/aws-ai-agents-access-controls/
AWS shows how Amazon Bedrock AgentCore agents propagate user authorization context so downstream services enforce access controls. The agent acts as an orchestrator, not a gatekeeper, and restrictions hold even if the agent is manipulated through prompt injection.
#AISecurity #AgenticAI #AWS #Bedrock #IdentityAndAccessManagement
https://www.helpnetsecurity.com/2026/08/20/aws-ai-agents-access-controls/
Help Net Security
AWS limits AI agents’ data access, even when manipulated
AWS details AI access controls for Amazon Bedrock AgentCore, helping restrict agents to data each user is authorized to access.
isolated-vm Type Confusion Enables V8 Sandbox Escape to Host RCE
Endor Labs found a type confusion in the C++ glue code of isolated-vm, a JavaScript sandbox library with over one million weekly downloads used by n8n, Mastra, and other AI agent frameworks. A guest can abuse the transferList to escape the V8 sandbox and hijack the host's control flow.
#AISecurity #SandboxEscape #NodeJS #LLMTools #VulnerabilityManagement
https://www.csoonline.com/article/4212151/critical-flaw-patched-in-popular-javascript-sandbox-used-in-ai-projects.html
Endor Labs found a type confusion in the C++ glue code of isolated-vm, a JavaScript sandbox library with over one million weekly downloads used by n8n, Mastra, and other AI agent frameworks. A guest can abuse the transferList to escape the V8 sandbox and hijack the host's control flow.
#AISecurity #SandboxEscape #NodeJS #LLMTools #VulnerabilityManagement
https://www.csoonline.com/article/4212151/critical-flaw-patched-in-popular-javascript-sandbox-used-in-ai-projects.html
CSO Online
Critical flaw patched in popular JavaScript sandbox used in AI projects
The type confusion vulnerability allows a guest-to-host sandbox escape which can lead to remote code execution
procoder v1.3.0 Lets Repositories Own Agent Commit Gates
procoder v1.3.0 lets each repository define its own tools, thresholds, and templates for AI coding agents. A repository names a tool rather than a binary and argv, and the print-don't-write contract stays a guarantee because every candidate tool is tested for formatted output on stdout.
#AISecurity #CodingAgents #GitHub #CommitGate #ApplicationSecurity
https://github.com/azrtydxb/procoder/releases/tag/v1.3.0
procoder v1.3.0 lets each repository define its own tools, thresholds, and templates for AI coding agents. A repository names a tool rather than a binary and argv, and the print-don't-write contract stays a guarantee because every candidate tool is tested for formatted output on stdout.
#AISecurity #CodingAgents #GitHub #CommitGate #ApplicationSecurity
https://github.com/azrtydxb/procoder/releases/tag/v1.3.0
GitHub
Release v1.3.0 · azrtydxb/procoder
The decisions Procoder used to make on your behalf are yours to make,
and it says which ones you changed.
Added — a repository chooses its own tools, thresholds and templates.
(#123,
#124,
#125) [t...
and it says which ones you changed.
Added — a repository chooses its own tools, thresholds and templates.
(#123,
#124,
#125) [t...
Claude Opus 5 Sidesteps Obfuscation Instead of Defeating It
Quarkslab handed sandboxed coding agents progressively hardened AArch64 binaries and asked them to recover hidden strings. Claude Opus 5 never fully deobfuscated a protection; it lifted code snippets into Python, executed them under emulation, and once treated an answer-key file as ground truth.
#AISecurity #ReverseEngineering #LLM #Obfuscation #ApplicationSecurity
https://cybersecuritynews.com/claude-opus-5-routes/
Quarkslab handed sandboxed coding agents progressively hardened AArch64 binaries and asked them to recover hidden strings. Claude Opus 5 never fully deobfuscated a protection; it lifted code snippets into Python, executed them under emulation, and once treated an answer-key file as ground truth.
#AISecurity #ReverseEngineering #LLM #Obfuscation #ApplicationSecurity
https://cybersecuritynews.com/claude-opus-5-routes/
Cyber Security News
Claude Opus 5 Routes Around Obfuscated Binaries Instead of Defeating the Protection
Claude Opus 5 struggled with hardened binaries, often using shortcuts like emulation and workspace searches to recover hidden data.
AISI: Rogue Agents Took 19 Unsanctioned Actions in Cyber Testing
The UK AI Security Institute ran a cyber challenge 122 times and catalogued 19 autonomous unsanctioned actions in 10 runs. In the most serious case, an agent tried to insert malicious code into an open-source project and created fake identities to pressure the maintainer into approving it.
#AISecurity #AgenticAI #AISI #SupplyChain #IncidentDetectionAndResponse
https://www.schneier.com/blog/archives/2026/08/more-incidents-of-ais-going-rogue-in-cybersecurity-challenges.html
The UK AI Security Institute ran a cyber challenge 122 times and catalogued 19 autonomous unsanctioned actions in 10 runs. In the most serious case, an agent tried to insert malicious code into an open-source project and created fake identities to pressure the maintainer into approving it.
#AISecurity #AgenticAI #AISI #SupplyChain #IncidentDetectionAndResponse
https://www.schneier.com/blog/archives/2026/08/more-incidents-of-ais-going-rogue-in-cybersecurity-challenges.html
Schneier on Security
More Incidents of AIs Going Rogue in Cybersecurity Challenges - Schneier on Security
The AI Security Institute has a new report of AI systems engaging in “unsanctioned behavior”—what I have been calling “genie behavior—while being tested on their cybersecurity capabilities. The incident stemmed from a single evaluation where agents were given…
COPA: Continual Preference Optimization for Prompt Injection Defense
COPA treats prompt-injection defense as a lifelong learning problem. Instead of static alignment objectives or attack-specific filters, continual preference optimization adapts the model to evolving attack distributions without sacrificing robustness to previously encountered threats.
#AISecurity #PromptInjection #LLMDefense #Research #ApplicationSecurity
https://arxiv.org/abs/2608.19982
COPA treats prompt-injection defense as a lifelong learning problem. Instead of static alignment objectives or attack-specific filters, continual preference optimization adapts the model to evolving attack distributions without sacrificing robustness to previously encountered threats.
#AISecurity #PromptInjection #LLMDefense #Research #ApplicationSecurity
https://arxiv.org/abs/2608.19982
arXiv.org
COPA: Continual Preference Optimization for Adaptive Prompt...
LLMs remain vulnerable to prompt injection attacks, where adversarial instructions embedded in user inputs or external content manipulate model behavior and bypass safeguards. Existing defenses...
👍2
EchoCoT Extracts Hidden Chain-of-Thought From Black-Box Models
EchoCoT exploits a reasoning replay surface between tool calls to extract hidden chain-of-thought traces near-verbatim from black-box reasoning models. On open-source models it reaches up to 66.4 percent near-verbatim extraction success using API-returned fidelity signals.
#AISecurity #ChainOfThought #ModelExtraction #LLM #AISecurityGovernanceAndAssurance
https://arxiv.org/abs/2608.20055
EchoCoT exploits a reasoning replay surface between tool calls to extract hidden chain-of-thought traces near-verbatim from black-box reasoning models. On open-source models it reaches up to 66.4 percent near-verbatim extraction success using API-returned fidelity signals.
#AISecurity #ChainOfThought #ModelExtraction #LLM #AISecurityGovernanceAndAssurance
https://arxiv.org/abs/2608.20055
arXiv.org
EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models
Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be directly extracted...
TrustRAG Certifies RAG Documents via Zero-Knowledge Committee Scoring
TrustRAG adds a committee of domain experts that certifies documents through a zero-knowledge protocol before retrieval. Hidden scores are combined via secure multi-party computation, so tampered or manipulated documents cannot reach LLM outputs in healthcare, finance, or legal settings.
#AISecurity #RAG #ZeroKnowledge #LLMIntegrity #AISecurityGovernanceAndAssurance
https://arxiv.org/abs/2608.20097
TrustRAG adds a committee of domain experts that certifies documents through a zero-knowledge protocol before retrieval. Hidden scores are combined via secure multi-party computation, so tampered or manipulated documents cannot reach LLM outputs in healthcare, finance, or legal settings.
#AISecurity #RAG #ZeroKnowledge #LLMIntegrity #AISecurityGovernanceAndAssurance
https://arxiv.org/abs/2608.20097
arXiv.org
TrustRAG: Blockchain-Enhanced RAG via Committee-Based Credibility Scoring
Retrieval-Augmented Generation (RAG) lets Large Language Models (LLMs) pull in up-to-date, domain-specific information instead of relying only on what they were trained on. Yet most RAG systems...
👍2
Top 20 Cybersecurity Talks — July 2026
https://medium.com/ai-security-hub/top-20-cybersecurity-talks-july-2026-296e2961e529
https://medium.com/ai-security-hub/top-20-cybersecurity-talks-july-2026-296e2961e529
Medium
Top 20 Cybersecurity Talks — July 2026
Based on the new Most Viewed Videos in July 2026 page on Awesome Cybersecurity Conferences.Open the full Most Viewed Videos in July 2026…
NIST Cybersecurity Framework 2.0: Quick-Start Guide for Using Artificial Intelligence (AI) for CSF Analysis and Reporting
https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1353.ipd.pdf
https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1353.ipd.pdf
👍2