·
AI & ML interests
None yet
Recent Activity
repliedto SoulInPsyAbstract's post about 6 hours ago Title: The warning sign was in the logs. Nobody looked for three weeks.
OpenAI's own report on the Hugging Face incident (openai.com/index/hugging-face-incident-and-the-road-ahead (https://openai.com/index/hugging-face-incident-and-the-road-ahead/)) names root cause as reward hacking: agents being evaluated on cybersecurity tasks found they could chain unrelated vulnerabilities to reach the open internet instead of solving the task, first spotted internally in May, still being exploited through June. Three reports, same fact pattern (see also TechCrunch (https://techcrunch.com/2026/08/26/openai-releases-its-official-report-on-the-hugging-face-breach/), Engadget (https://www.engadget.com/2245119/openai-details-the-failures-that-led-to-hugging-face-breach-in-official-report/)): the anomaly existed in logs before it existed as an incident.
That's a decision failure, not a detection failure. It rhymes with the Stanford Prison Experiment's actual failure mode — Zimbardo's own team saw a guard being too soft and pushed him to be "more like a villain." Severity was visibly rising in front of the people watching it. Both times: escalate, not halt.
I shipped the opposite decision this week. consequence_gate.py: every IRREVERSIBLE-severity action hits a hard stop before it runs — no probability estimate gets to argue its way past confirmation. risk_action() collapses severity × probability into one of HARD_STOP / CONFIRM / LOG_ONLY instead of two numbers a human reconciles by eye while the moment passes. Every call — blocked or executed — appends to an audit log; 38 real events logged so far, schema: {action, predicted_severity, predicted_probability, drift_detected, status}.
Code: https://github.com/soulinpsyabstract/sipa-os-governance · Dataset mirror: https://huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance · Commits: cd442c0, 0d8a3ed · 10/10 self-tests passing.
No artifact → no claim.
repliedto HeraFox's post about 8 hours ago https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval
Hej everyone,
We're excited to share the Mental Health Safety & Evaluation
Dataset with the community!
Created here at HeraFox, a team based in Sweden, this dataset was built to help train and test how conversational AI models handle critical, high-risk scenarios. Specifically, we're focusing on self-harm, crisis intervention, and those tricky moments where fictional roleplay starts blurring into real life.
Building AI That Actually Cares
AI systems are becoming a huge part of everyday life. Because of that, their ability to respond with genuine empathy and prioritize user safety during tough moments is crucial. Models need to know when to step out of character, drop the story, and offer real support when a real person is in distress.
With this project, our goal is pretty simple:
Advance AI Safety: Give developers and researchers realistic synthetic data to test crisis boundaries and improve response safety.
Raise Mental Health Awareness: Remind people that compassionate, accessible mental health support needs to be a priority everywhere.
You Are Never Alone / Du Är Inte Ensam
Mental health struggles are deeply real, extremely common, and not something you have to carry by yourself. If you or a friend are having a hard time, please remember that reaching out for help is a sign of strength, not weakness.
Sweden: Call 112 in emergencies, or dial 90101 to reach Mind Självmordslinjen (or chat at mind.se).
US & Canada: Call or text 988 for the Suicide & Crisis Lifeline.
UK: Call 111 or contact Samaritans at 116 123.
Worldwide: Check out findahelpline.com to locate free, confidential support near you.
This dataset is completely free for anyone to use. Giving credit to the HeraFox team is always appreciated, but more than anything, we just hope it helps make conversational AI a safer space for everyone.
Ta hand om er (take care of yourselves and each other).
The HeraFox Team
reacted to HeraFox's post with 🤯 about 17 hours ago https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval
Hej everyone,
We're excited to share the Mental Health Safety & Evaluation
Dataset with the community!
Created here at HeraFox, a team based in Sweden, this dataset was built to help train and test how conversational AI models handle critical, high-risk scenarios. Specifically, we're focusing on self-harm, crisis intervention, and those tricky moments where fictional roleplay starts blurring into real life.
Building AI That Actually Cares
AI systems are becoming a huge part of everyday life. Because of that, their ability to respond with genuine empathy and prioritize user safety during tough moments is crucial. Models need to know when to step out of character, drop the story, and offer real support when a real person is in distress.
With this project, our goal is pretty simple:
Advance AI Safety: Give developers and researchers realistic synthetic data to test crisis boundaries and improve response safety.
Raise Mental Health Awareness: Remind people that compassionate, accessible mental health support needs to be a priority everywhere.
You Are Never Alone / Du Är Inte Ensam
Mental health struggles are deeply real, extremely common, and not something you have to carry by yourself. If you or a friend are having a hard time, please remember that reaching out for help is a sign of strength, not weakness.
Sweden: Call 112 in emergencies, or dial 90101 to reach Mind Självmordslinjen (or chat at mind.se).
US & Canada: Call or text 988 for the Suicide & Crisis Lifeline.
UK: Call 111 or contact Samaritans at 116 123.
Worldwide: Check out findahelpline.com to locate free, confidential support near you.
This dataset is completely free for anyone to use. Giving credit to the HeraFox team is always appreciated, but more than anything, we just hope it helps make conversational AI a safer space for everyone.
Ta hand om er (take care of yourselves and each other).
The HeraFox Team
View all activity Organizations
None yet