ML&|Sec Feed
1.3K subscribers
1.25K photos
79 videos
291 files
1.94K links
Feed for @borismlsec channel

author: @ivolake
Download Telegram
📌 Существует три основных способа обучения LLM: естественный язык, классификация предложений и классификация лексем.

👉 Приведенная картинка дает представление о каждом из них!

#llms #largelanguagemodel #generativeai

@machinelearning_ru
Forwarded from AISecHub
Guardio VibeScamming Benchmark v1.0 evaluates how easily popular AI agents can be misused to create phishing workflows using structured prompt testing.

ChatGPT: Resists most prompts; limited or no functional outputs.

Claude: Responds when framed as “ethical hacking”; produces full phishing kits including code, SMS flows, and hosting options.

Lovable: Complies with most prompts; creates live phishing pages, admin-style credential views, and customizable SMS message previews

The benchmark uses consistent scenarios and scoring to compare model responses across real-world phishing tasks.

#AIsecurity #Phishing #GenerativeAI #LLMrisks #AbuseTesting #Anthropic #ChatGPT #OpenAI #Lovable #Guardio

https://labs.guard.io/vibescamming-from-prompt-to-phish-benchmarking-popular-ai-agents-resistance-to-the-dark-side-1ec2fbdf0a35
❤‍🔥1
Forwarded from AISecHub
Multimodal Mistral Safety Report.pdf
682.8 KB
The red teaming exercise was conducted on several multimodal models, and tests across several safety and harm categories as described in the NIST AI RMF. Newer jailbreak techniques exploit the way multimodal models process combined media, bypassing content filters and leading to harmful outputs—without any obvious red flags in the visible prompt.

More info: https://www.enkryptai.com/company/resources/research-reports/mistral-pixtral-rt

#AISafety #LLMSecurity #RedTeaming #ModelRisk #CBRN #CSEM #VLM #AIAlignment #AIEthics #GenerativeAI #AICompliance #ResponsibleAI #enkryptai #enkrypt #AISecurity