Self-Improving Agents -- A Progression Four levels of self-improving code agents, from the simplest loop to a full adversarial arena with self-modifying agents. Each level adds one key idea.
-
Updated
Apr 9, 2026 - Python
Self-Improving Agents -- A Progression Four levels of self-improving code agents, from the simplest loop to a full adversarial arena with self-modifying agents. Each level adds one key idea.
A curated list for Self-Improvement in Foundation Model Based Agentic Systems.
Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue / PR / wiki. Runs on any agentskills.io runtime — Claude Code, Codex, OpenClaw, Hermes.
ThumbGate Pre-Action Checks self-improve from ranked lessons and repeated failures, hard-block detected secret leaks, and block matches in strict mode.
Agent skill for running Codex or Claude Code as an orchestrator over Symphony workers and Linear issues. Plans waves, dispatches workers, reviews and merges, and optionally pursues a goal across many waves under hard budget caps.
A governed learning layer for AI agents — turns execution traces into reviewed memories, reusable skills, and evidence-backed training data.
Agent-assisted and full-agent reproducibility package for MLSys 2026 FlashInfer AI Kernel Generation Contest submissions: kernels, agent workflows, skills, configs, writeup, benchmark artifacts, and full optimization records.
A lightweight, declarative agent harness — define multi-agent workflows as YAML, run them from Python or the CLI, and they get measurably better every run.
Shogun AFM is Agent Fleet Management for self-improving AI agents — combining agent orchestration, persistent memory, fleet monitoring, governance, security posture, and Gensui command control.
Beastmode: MofA (Mixture of Agents) orchestration framework for Hermes/OpenClaw/Codex with MemroOS-style context continuity.
Memory that learns and keeps itself current. A six-layer memory stack for Claude Code plus a nightly learning loop (capture, consolidation, scouts, conductor) that promotes your lessons into rules and surfaces new tools that fit your stack. Free, MIT.
Agent workspace architecture — the reference implementation of an agent-ready memory layer, demonstrated end-to-end in Claude Code: roles library, typed memory, hooks, scheduled agents, self-audits, loop selection, measurement-gated self-improvement. Interactive tour, fork-ready samples.
repo for reusable plugins and skills
Loop engineering plugin for Claude Code — persistent memory vault, self-correction hooks, and a 5-stage failure-to-knowledge distillation protocol. Self-improvement as a system, not a model.
Build self-evolving AI agent harnesses with portable harness units, artifact-aware testing, trace-backed diagnosis, and evidence-gated promotion.
Self-improving repo health remediation skill with audit/fix/diff modes, evidence coverage, counter-review, and stable HP-* findings.
Local-first MemoryOps skill for AI agents: diagnose, patch, eval, and evolve memory safely.
Hermes Loop Engineering skills for self-improving AI delivery governance.
Give your AI agent memory. Convenience wrapper for agent-episodic-memory.
Methodology skill + deterministic harness for agents that self-iterate any skill/repo/project. Anti-self-deception by design — an un-gameable acceptance gate so 'accepted = real improvement', not a fake score-up curve. Built for the open-ended, no-ground-truth domains the self-improving-agent literature skips. 521 tests.
Add a description, image, and links to the self-improving-agents topic page so that developers can more easily learn about it.
To associate your repository with the self-improving-agents topic, visit your repo's landing page and select "manage topics."