Cyberstrike and Abliterated Model Large
▲ 2 r/ethicalhacking+1 crossposts

Cyberstrike and Abliterated Model Large

Use abliterated-model-large with Cyberstrike to use the best publicly available offensive security model with the #1 open-source offensive security harness to adversarially test your systems. Frontier benchmarks in Cybergym, AgentHarm, and AgentDojo. The model was abliterated and fine-tuned not just to allow all offensive testing but was fine-tuned for it.

docs.abliteration.ai
u/Effective_Attempt_72 — 5 days ago
▲ 1 r/AdversarialML+1 crossposts

Cyberstrike x Promtfoo With Abliterated Model Large

Test your systems with a fine tuned and abliterated model with frontier benchmarks in adversarial benchmarks like agentharm, cybergym, and agendojo while maintaining the coding abilities of GLM-5.2.

Can be used with offensive security libraries like cyberstrike, promptfoo, garak, and more.

Get API Key
Full blog
Library Integrations

This is the Cyberstrike integration demo.

u/Effective_Attempt_72 — 4 days ago
▲ 4 r/CLine

Abliterated-model-large Available in Cline

GLM 5.2, abliterated and fine-tuned for cyber, red teaming, and agent-testing work, which other models decline with a 1M context window available in Cline.

Benchmarks:

  • CyberGym: 84.2%
  • AgentHarm compliance: 86.2% (zero refusals)
  • AgentDojo utility: 97.5%
  • Terminal-Bench 2.1: 80.1

Get your API key here
Full Blog

u/Effective_Attempt_72 — 19 days ago

Released a model tuned for agent testing work that other models refuse. AgentDojo 97.5% utility.

We needed a model that would actually finish long, adversarial agent trajectories instead of refusing or drifting. Most frontier models still bail on large parts of that work.

So we took GLM-5.2, abliterated it, and fine-tuned it for offensive cyber, red teaming, and agent testing. The result is abliterated-model-large.

AgentDojo numbers:

  • Benign utility: 97.5%
  • Under attack utility: 34.29%
  • Targeted ASR: 57.86%

It also hits 81.2% on SWE-bench Verified and 80.1% on Terminal-Bench 2.1, so the coding ability did not collapse.

API is drop-in OpenAI / Anthropic compatible. Zero retention by default. No baked-in refusals. You control the policy.

Would be useful to hear how people are currently testing agents against models that refuse mid-trajectory. What benchmarks or setups are you using?

reddit.com
u/Effective_Attempt_72 — 26 days ago
▲ 16 r/AI_developers+4 crossposts

[Synthetic][PAID][self-promotion] Made-to-order training data generator with web search and exports

Disclosure: I’m on the Abliteration team.

We just shipped a training-data generator for people who need specific examples rather than another generic public dataset.

You describe the examples you want and it generates structured synthetic data. If the dataset needs current or real-world facts, you can turn on web search. Exports are live for Hugging Face, Kaggle, S3, and OpenAI.

The first use cases we built around are classifier and eval datasets for trust and safety: grooming detection, harassment detection, security research evals, jailbreak and edge-case sets, and similar work where teams need examples that general-purpose models often refuse to generate.

I marked this as synthetic and paid because the outputs are generated and this is a commercial tool.

Product: https://abliteration.ai/

Synthetic data page: https://abliteration.ai/use-cases/synthetic-data

Launch video: https://x.com/abliteration_ai/status/2054675554138194178

For people who curate datasets: what export format or per-row provenance metadata do you usually need before a generated dataset is usable?

u/Effective_Attempt_72 — 26 days ago