A curated list for Self-Improvement in Foundation Model Based Agentic Systems.
-
Updated
Aug 24, 2026 - TeX
A curated list for Self-Improvement in Foundation Model Based Agentic Systems.
Self-Improving Agents -- A Progression Four levels of self-improving code agents, from the simplest loop to a full adversarial arena with self-modifying agents. Each level adds one key idea.
Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue / PR / wiki. Runs on any agentskills.io runtime — Claude Code, Codex, OpenClaw, Hermes.
A curated research map of Recursive Self-Improvement (RSI): models, agents, harnesses, embodied systems, automated AI R&D, benchmarks, and safety.
A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and auditable.
ThumbGate Pre-Action Checks self-improve from ranked lessons and repeated failures, hard-block detected secret leaks, and block matches in strict mode.
The Dream Machine — a config-driven engine for nightly, cloud-scheduled, evidence-gated repository evolution. Composes @metaharness/flywheel, darwin, and redblue behind a promotion gate that never merges.
Agent-assisted and full-agent reproducibility package for MLSys 2026 FlashInfer AI Kernel Generation Contest submissions: kernels, agent workflows, skills, configs, writeup, benchmark artifacts, and full optimization records.
Beastmode: MofA (Mixture of Agents) orchestration framework for Hermes/OpenClaw/Codex with MemroOS-style context continuity.
Agent skill for running Codex or Claude Code as an orchestrator over Symphony workers and Linear issues. Plans waves, dispatches workers, reviews and merges, and optionally pursues a goal across many waves under hard budget caps.
Agent swarms that recursively evolve themselves — recursive self-improvement as auditable topology. Build an agentic harness swarm as a directory tree. One Rust binary.
A governed learning layer for AI agents — turns execution traces into reviewed memories, reusable skills, and evidence-backed training data.
Agent-ready knowledge architecture, run daily in Claude Code: turn a coding agent into a system you can hand work to and trust while you're away. Roles library, typed memory, hooks, scheduled agents, delegation queue, self-audits, token budgeting, measurement-gated self-improvement. Interactive tour, fork-ready samples.
A curated, evidence-aware collection of recursive self-improvement research, agents, harnesses, benchmarks, and safety work.
A lightweight, declarative agent harness — define multi-agent workflows as YAML, run them from Python or the CLI, and they get measurably better every run.
DeepSeek Harness (DSH) plugin for self-improving AI agents: continual learning, persistent memory, cross-session knowledge, review-and-refine workflows, and automatic rollback.
Self-improving agents, governed. Areev is the substrate for adaptive agents — agents that get better from their own history, under human authority, in steps you can inspect, undo, and re-measure.
Shogun AFM is Agent Fleet Management for self-improving AI agents — combining agent orchestration, persistent memory, fleet monitoring, governance, security posture, and Gensui command control.
repo for reusable plugins and skills
Memory that learns and keeps itself current. A six-layer memory stack for Claude Code plus a nightly learning loop (capture, consolidation, scouts, conductor) that promotes your lessons into rules and surfaces new tools that fit your stack. Free, MIT.
Add a description, image, and links to the self-improving-agents topic page so that developers can more easily learn about it.
To associate your repository with the self-improving-agents topic, visit your repo's landing page and select "manage topics."