Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes.
Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench.
I've spent a lot of my free time lately to do something similar to this (except not with Qwen), because I was dismayed by the error rates and compute cost of taking a published skill and finding LLMs stuck in loops for hours doing "oh this doesn't work, let me try xyz". So instead I read the traces, pass them into a little framework I built for finding errors / hot spots, try to find root causes (with or without bot help) and tune skill files, prompts, harness config... and so on. It doesn't really matter which model you use either, as long as you make sure that your improved prompts are consistent, i.e. you can make GLM-5.2 as good as Kimi K3, simply by taking away the causes of confusion.
Bottom line, LLM integration tuning is to me still the more valuable exercise, and looping through traces is currently the best way to do that.
Are you using multiple harnesses like this or doing recursive rewrites with a single harness?
I use multiple in the first round, but since they are doing RL on the model as the next step and I am doing rewrites of everything other than the model, I generally just standardize it back to pi-agent because that is the easiest to integrate/customize.
For example, one of the things Claude Code does well is subagent creation, whereas on pi, with the
pi-subagentsplugin, I was wasting tons of tokens on default profiles that are useless and each run the bots finding out that these are useless. So I tuned that by deleting all the standard agent profiles, removing all the docs and tools that were no longer relevant, and just made a clear instruction with a singular path to success in creating a subagent.It's a lot cheaper to tune the skill file and plugins than it is to finetune a model. Trying to not boil oceans, haha.
This is how autonomous business agents will actually work in 2027. Not bigger models, but RSR - discover with 3 harnesses, rewrite into 1 clean runbook.
Same as Aba market: 3 apprentices try different ways to fix phone, master watches their traces, writes one best procedure manual. Now every new apprentice runs the manual and solves 34.3% more.
Bitcoin angle: Qwen-3.8-27B = cheap base money, runbooks = Lightning invoices, critic = verification. You don't need GPT-5 to run a business fleet, you need self-custody of your successful trajectories.
Future: every PH logistics company will have this - planner, critic, executor loop. That 11k rewritten trajectories is your real moat.