Skip to content

Phase 5 — Counterfactual Replay, Interaction Attribution... - #6

Merged
DSCmatter merged 4 commits into
mainfrom
phase5
Sep 6, 2026
Merged

DSCmatter merged 4 commits into
mainfrom
phase5

Conversation

@Garry400

@Garry400 Garry400 commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

CodeAnt-AI Description

Add safe counterfactual replay, interaction analysis, and minimal causal slices

What Changed

  • Decisions can be replayed with selected inputs replaced by baseline values to show whether the outcome changes.
  • Replay refuses decisions that could trigger payments, emails, database writes, or external mutations.
  • Added replay, interaction, and minimize CLI commands for comparing outcomes, measuring individual and joint input effects, and reducing explanations to the smallest failure-reproducing event set.
  • Decision evaluators can be registered by decision type, including the fixture’s customer approval policy.
  • Invalid decision IDs, intervention ports, duplicate ports, and sample counts now produce clear errors instead of incorrect or unstable analysis.
  • Documentation and tests cover replay safety, outcome changes, interaction results, minimal slicing, caching, budget limits, and CLI output.

Impact

✅ Safer decision analysis without triggering side effects
✅ Clearer causes of policy failures
✅ Four-event minimal explanations for the fixture failure

💡 Usage Guide

Checking Your Pull Request

Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.

Talking to CodeAnt AI

Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:

@codeant-ai ask: Your question here

This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.

Example

@codeant-ai ask: Can you suggest a safer alternative to storing this secret?

Preserve Org Learnings with CodeAnt

You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:

@codeant-ai: Your feedback here

This helps CodeAnt AI learn and adapt to your team's coding style and standards.

Example

@codeant-ai: Do not flag unused imports.

Retrigger review

Ask CodeAnt AI to review the PR again, by typing:

@codeant-ai: review

Check Your Repository Health

To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.

@codeant-ai

codeant-ai Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

🤖 CodeAnt AI — Review Status

Status Commit Started (UTC) Finished (UTC)
✅ Incremental review completed 9f8ead2 Sep 06, 2026 · 05:33 05:33
✅ Reviewed your PR 1f36c8c Sep 03, 2026 · 18:30 18:33

@codeant-ai

codeant-ai Bot commented Sep 3, 2026

Copy link
Copy Markdown

Thanks for using CodeAnt! 🎉

We're free for open-source projects. if you're enjoying it, help us grow by sharing.

Share on X ·
Reddit ·
LinkedIn

@codeant-ai codeant-ai Bot added the size:XL This PR changes 500-999 lines, ignoring generated files label Sep 3, 2026
Comment thread cli/main.py Outdated
Comment thread core/replay.py Outdated
Comment thread core/replay.py
Comment thread core/replay.py Outdated
Comment on lines +115 to +121
downstream_event_id = contract.metadata.get("downstream_failure_event")
if downstream_event_id:
getter = getattr(log, "get", None)
if callable(getter):
downstream_ev = getter(downstream_event_id)
if downstream_ev is not None:
reconstruct(downstream_ev.agent_id, downstream_ev.logical_seq, log=log)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: Downstream replay reconstructs the recorded state but discards it, so counterfactual differences are never compared or reported. [api mismatch]

Assessment: 🟠 Major · 🔁 Occurrence: Sometimes

Use CodeAnt Skill Fix in Cursor Fix in VSCode Claude

Prompt for AI Agent 🤖
This is a comment left during a code review.

**Path:** core/replay.py
**Line:** 115:121
**Comment:**
	*Api Mismatch: Downstream replay reconstructs the recorded state but discards it, so counterfactual differences are never compared or reported.

Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix
👍 | 👎

Comment thread core/replay.py
Comment thread core/replay.py
@DSCmatter DSCmatter added documentation Improvements or additions to documentation enhancement New feature or request labels Sep 3, 2026
- cli/main.py: remove silent fallback to fixture for unknown decision IDs;
  unknown IDs now raise ValueError instead of silently analyzing A3
- core/replay.py: guard metadata.get() against None for deserialized contracts
  with null metadata field
- core/replay.py: validate intervention port IDs against contract.ports,
  raising ValueError for unknown/misspelled port IDs
- core/replay.py: detect duplicate port IDs before compute_shapley_interaction
  to prevent corrupted Shapley attribution results
- core/replay.py: validate samples_per_cell >= 1 to prevent ZeroDivisionError
  in calc_v_from_cells

Not resolved: downstream replay state diff reporting (deferred, documented)
@DSCmatter

Copy link
Copy Markdown
Contributor

@CodeAnt-AI review this again with the new PR

@codeant-ai

codeant-ai Bot commented Sep 3, 2026

Copy link
Copy Markdown

Question: review this again with the new pr

Answer:

Findings

1. replay/interaction can silently use the bundled fixture for PostgreSQL events

Location: cli/main.py, resolve_decision_contract

if decision_or_event_id == "dec_customer_approval_A3" or decision_or_event_id in ("A3", "A4"):
    return create_fixture_decision(fixture_data)

When the CLI is run without --fixture, fixture_data is None, so create_fixture_decision() loads the repository’s local fixture/fixture.json. This means:

uv run python -m cli.main replay A3 ...

can replay the bundled fixture rather than the decision represented by the PostgreSQL event being queried. The command may therefore report a successful-looking result for the wrong run or tenant.

Also, A4 is a downstream event, not the decision event itself, so resolving it directly to the fixture decision is an implicit special case that does not generalize to other decisions.

Suggestion: Resolve the event from log first and extract its embedded DecisionContract; only use create_fixture_decision() when the fixture backend explicitly identifies the canonical fixture decision. For example, avoid hardcoded IDs and return an error if an event does not contain a contract.


2. Downstream replay does not actually validate the counterfactual downstream result

Location: core/replay.py, counterfactual_replay

if downstream_event_id:
    ...
    if downstream_ev is not None:
        reconstruct(downstream_ev.agent_id, downstream_ev.logical_seq, log=log)
return decision_outcome

mode="downstream_replay" is documented as validating downstream events against recorded log states, but the implementation only reconstructs the recorded state and discards it. It does not compare the reconstructed state or downstream outcome with the counterfactual decision result.

Consequently, any evaluator result is returned as valid as long as reconstruction does not raise, even if the counterfactual would produce a different downstream state or failure. This makes the mode equivalent to recorded_output plus an unrelated integrity read.

Suggestion: Define the expected downstream invariant and compare the reconstructed/counterfactual downstream result against it. If downstream replay is not implemented yet, raise NotImplementedError rather than claiming reconciliation.


3. Side-effect safety is only enforced when the contract metadata is already trusted

Location: core/replay.py, counterfactual_replay

The safety check only examines:

contract.is_side_effecting
contract.metadata["tool_category"]

It does not inspect the decision’s associated events, tool calls, or downstream operations. A contract with is_side_effecting=False and no tool_category can still represent a decision that sends an email, performs a payment, writes to a database, or mutates an external API.

This does not satisfy the stated guarantee that replay refuses decisions that could trigger those side effects; it only refuses contracts that have been correctly annotated beforehand.

Suggestion: Make the contract creation/loading path derive or verify side-effect metadata from the associated event/tool records, and fail closed when side-effect status is unknown. At minimum, use an explicit “sandboxed/read-only” capability rather than treating a missing flag as safe.


4. The interaction “p-value” is not a statistically valid significance test

Location: core/replay.py, bootstrap packaging

non_pos = sum(1 for x in vals if x <= 0.0)
p_val = non_pos / len(vals)

This is the fraction of bootstrap estimates that are non-positive, not a calibrated p-value. It also has no correction for the many pairwise hypotheses being tested. With deterministic binary cells, the bootstrap samples are duplicated observations rather than independent model trials, so the resulting confidence/significance values can be misleading.

The README and CLI describe these values as confidence/significance results, and the tests assert p_value < 0.01, so this is user-visible rather than merely an implementation detail.

Suggestion: Either label this as a bootstrap sign proportion (and avoid calling it a p-value), or implement a defined hypothesis test/permutation test with an appropriate null distribution and document the assumptions.

@DSCmatter

Copy link
Copy Markdown
Contributor

@Garry400 fix this too

@codeant-ai codeant-ai Bot added size:XXL This PR changes 1000+ lines, ignoring generated files and removed size:XL This PR changes 500-999 lines, ignoring generated files labels Sep 6, 2026
@DSCmatter
DSCmatter merged commit b3121fd into main Sep 6, 2026
2 checks passed
@DSCmatter
DSCmatter deleted the phase5 branch September 6, 2026 05:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation enhancement New feature or request size:XXL This PR changes 1000+ lines, ignoring generated files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants