Skip to content

Consider whether the "adversarial review" framing can be softened #760

Description

@lunaynx

I had a Fable thread where I essentially looped /codex:adversarial-review -> ask Claude to fix -> repeat. The work was not cybersecurity related at all, it only involved simple networking. And yet, on about the 6th cycle (yeah, it's cooked, don't ask), the "Verify and address Codex's findings" prompt tripped Fable's cyber classifier.

Now, I obviously can't know for sure whether this is what tripped it. But I wouldn't be surprised at all if having "adversarial review" repeated 6 times in the context (and likely double that if you include the CoT), would trip the classifier. The term "adversarial" has a cybersecurity/red teaming connotation, even though it isn't being used that way there. So I wonder if maybe it's worth replacing it with a different term to reduce the risk of false flags with frontier models.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions