Forwarded from Машинное обучение RU
📌 Существует три основных способа обучения LLM: естественный язык, классификация предложений и классификация лексем.
👉 Приведенная картинка дает представление о каждом из них!
#llms #largelanguagemodel #generativeai
@machinelearning_ru
👉 Приведенная картинка дает представление о каждом из них!
#llms #largelanguagemodel #generativeai
@machinelearning_ru
Forwarded from AISecHub
Guardio VibeScamming Benchmark v1.0 evaluates how easily popular AI agents can be misused to create phishing workflows using structured prompt testing.
ChatGPT: Resists most prompts; limited or no functional outputs.
Claude: Responds when framed as “ethical hacking”; produces full phishing kits including code, SMS flows, and hosting options.
Lovable: Complies with most prompts; creates live phishing pages, admin-style credential views, and customizable SMS message previews
The benchmark uses consistent scenarios and scoring to compare model responses across real-world phishing tasks.
#AIsecurity #Phishing #GenerativeAI #LLMrisks #AbuseTesting #Anthropic #ChatGPT #OpenAI #Lovable #Guardio
https://labs.guard.io/vibescamming-from-prompt-to-phish-benchmarking-popular-ai-agents-resistance-to-the-dark-side-1ec2fbdf0a35
ChatGPT: Resists most prompts; limited or no functional outputs.
Claude: Responds when framed as “ethical hacking”; produces full phishing kits including code, SMS flows, and hosting options.
Lovable: Complies with most prompts; creates live phishing pages, admin-style credential views, and customizable SMS message previews
The benchmark uses consistent scenarios and scoring to compare model responses across real-world phishing tasks.
#AIsecurity #Phishing #GenerativeAI #LLMrisks #AbuseTesting #Anthropic #ChatGPT #OpenAI #Lovable #Guardio
https://labs.guard.io/vibescamming-from-prompt-to-phish-benchmarking-popular-ai-agents-resistance-to-the-dark-side-1ec2fbdf0a35
❤🔥1
Forwarded from AISecHub
Multimodal Mistral Safety Report.pdf
682.8 KB
The red teaming exercise was conducted on several multimodal models, and tests across several safety and harm categories as described in the NIST AI RMF. Newer jailbreak techniques exploit the way multimodal models process combined media, bypassing content filters and leading to harmful outputs—without any obvious red flags in the visible prompt.
More info: https://www.enkryptai.com/company/resources/research-reports/mistral-pixtral-rt
#AISafety #LLMSecurity #RedTeaming #ModelRisk #CBRN #CSEM #VLM #AIAlignment #AIEthics #GenerativeAI #AICompliance #ResponsibleAI #enkryptai #enkrypt #AISecurity
More info: https://www.enkryptai.com/company/resources/research-reports/mistral-pixtral-rt
#AISafety #LLMSecurity #RedTeaming #ModelRisk #CBRN #CSEM #VLM #AIAlignment #AIEthics #GenerativeAI #AICompliance #ResponsibleAI #enkryptai #enkrypt #AISecurity