Self-Collaboration Code Generation via ChatGPT: How It Works
Self-collaboration code generation via ChatGPT means organizing AI-assisted programming into separate responsibilities, with explicit feedback between them. Instead of accepting the first implementation, you establish the requirements, produce a candidate solution, examine it against those requirements, and repair specific failures.
You can try a manual version through a sequence of ChatGPT prompts. An automated version requires an environment that actually coordinates agent runs and passes their outputs between stages. Writing three role names in a message does not, by itself, establish that three separate agents were launched.

What the Original Self-Collaboration Research Proposed
The phrase is also the title of a paper by Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li, first submitted in April 2023 and revised in May 2024. Their framework assigned analyst, coder, and tester roles to language-model agents, using shared information and feedback to refine code.
One detail matters when applying the idea: the paper's tester simulated testing and produced a report rather than executing code during that stage. The researchers separately evaluated generated solutions using benchmark tests.
The study reported improvements under its experimental conditions, including experiments with the historical gpt-3.5-turbo-0301 model. Those findings do not establish a guaranteed improvement for today's models, your repository, or a copied prompt.
The practical workflow below is my adaptation for everyday development. It adds explicit acceptance checks and real execution where a suitable runtime is available; it is not a reproduction of the paper's experiment.
Choose How You Will Run the Workflow
Before choosing elaborate prompts, decide where the work and verification will happen.
| Approach | How work moves | Main trade-off |
|---|---|---|
| One conversation, successive prompts | You request planning, implementation, and review in separate turns. | Simple to manage, but the review sees the earlier discussion. |
| Separate implementation and review conversations | You provide the reviewer with the requirements and current code. | A different conversational starting point, with manual context transfer. |
| Automated agent workflow | Software coordinates agents, tools, artifacts, and stopping conditions. | More control at scale, with additional setup and coordination costs. |
A separate chat does not guarantee an independent judgment. If both reviewers receive the same mistaken requirement, both may produce a plausible answer to the wrong problem. Give the reviewer the original task and current implementation, and ask for evidence behind each finding.
OpenAI documents genuine multi-agent execution through its Responses API, including a coordinating root agent and subagents. Its Codex and Agents SDK cookbook also demonstrates staged handoffs between specialist agents. These are implementation options for automation, rather than prerequisites for trying a manual workflow.
Start with an Acceptance Contract
An acceptance contract is a short description of the behavior the result must satisfy. Write it before generating code so that implementation choices do not quietly redefine success.
For a real project, supply the relevant files or code, the language and framework versions, the existing interface, and the checks the project already uses. Official OpenAI review guidance similarly recommends identifying the files or revision to inspect and the criteria for the review.
I would include these five items:
- Outcome: the observable behavior you want.
- Inputs and outputs: accepted values, return types, and failure behavior.
- Compatibility: interfaces and existing behavior that must remain stable.
- Scope: the files and dependencies the task may change.
- Verification: examples, existing tests, or commands that can demonstrate success.
For example, “clean up my article titles” leaves several decisions open. Should duplicates be case-sensitive? Should blank entries disappear? Should the original order survive? Resolving these questions is more valuable than asking three agents to debate an ambiguous sentence.
A Practical ChatGPT Coding Workflow
1. Ask for a brief implementation plan
Provide the acceptance contract and ask ChatGPT to identify the relevant interface, likely changes, and unresolved decisions. Request a concise plan rather than a lengthy fictional meeting.
For a small utility, the plan might be only four sentences. A repository change may require inspecting callers and existing tests first. Let the assistant gather enough information to understand the task, while keeping the proposed patch focused.
Review decisions that alter intended behavior before implementation. If the specification is already complete, there is no benefit in adding an approval ceremony to every trivial step.
2. Generate one bounded candidate
Ask for code that follows the agreed contract and fits the existing project. Specify whether you want a complete function, an updated file, or a patch.
I would keep unrelated cleanup out of the first candidate because it makes failures harder to attribute. If a necessary change extends beyond the original scope, request an explanation of that dependency rather than an unexplained rewrite.
When scope expansion is a recurring problem, the guide to unrelated file changes explains how inspection, editing, and automated commands can affect different parts of a repository.
3. Review against the original requirements
Give the reviewer the current candidate and the acceptance contract. Ask it to find a specific counterexample: an input or execution path where the result violates a requirement.
A useful finding contains the violated rule, the triggering condition, and the expected behavior. “Improve robustness” is too vague to justify a change. “An empty title survives even though blank entries must be removed” gives the next stage a concrete problem to solve.
Separate confirmed defects, questions requiring execution, and optional style suggestions. The implementer should not have to guess which comments actually block acceptance.
4. Execute the checks that matter
If the available environment can run the project, request the appropriate tests and inspect their output. If it cannot, run the supplied checks locally and return the actual error or failure message.
Keep these states distinct: a test was proposed, a test was executed, or a test passed. Only the last two require execution evidence. A natural-language walkthrough of expected behavior does not establish a successful test run.
Also verify the runtime matches the task. Running an isolated function does not demonstrate that a web application builds, that its dependency versions are compatible, or that a browser interaction works.
5. Repair specific failures and finish
Return the failure details to the implementation stage. Ask for the smallest complete correction, then repeat the affected checks.
For a small task, I would set a modest revision limit, such as two repair rounds. That is a workflow choice, not an optimal number established by research. If the same failure survives, stop adding similar prompts and investigate missing context, an incorrect contract, or an environment mismatch.
Finish when the agreed checks pass and the change satisfies the requirements, with any remaining limitation stated. Another stylistic opinion is not automatically a reason to continue editing.
Worked Example: Removing Duplicate Article Titles
Consider an illustrative Python utility for a content workflow. It accepts a list of strings and returns cleaned titles with duplicates removed.
The contract is specific: trim surrounding whitespace, ignore blank entries, compare titles using Unicode case folding, preserve the first occurrence's trimmed spelling, and preserve input order. It uses only the standard library.
A shortcut such as list(set(titles)) does not meet that contract. It does not perform the requested cleaning or case-insensitive comparison, and it does not guarantee the required order.
The review should turn those requirements into examples:
| Input | Expected result |
|---|---|
[" AI Tools ", "ai tools", "Python"] | ["AI Tools", "Python"] |
["", " ", "News"] | ["News"] |
["B", "A", "b"] | ["B", "A"] |
[] | [] |
A candidate matching those stated rules is:
def unique_titles(titles):
result = []
seen = set()
for title in titles:
cleaned = title.strip()
key = cleaned.casefold()
if key and key not in seen:
seen.add(key)
result.append(cleaned)
return result
assert unique_titles([" AI Tools ", "ai tools", "Python"]) == ["AI Tools", "Python"]
assert unique_titles(["", " ", "News"]) == ["News"]
assert unique_titles(["B", "A", "b"]) == ["B", "A"]
assert unique_titles([]) == []
assert unique_titles(["Straße", "STRASSE"]) == ["Straße"]Save the example as a Python file and run it with your Python interpreter. A failing assertion raises an error; the assertions produce no output when they succeed. These checks cover the listed behavior, rather than proving correctness for every possible requirement.
The final example also makes the case-folding requirement observable. If your publication wants different treatment of Unicode titles, change the contract first. A reviewer should not silently substitute its own editorial policy.
A Prompt You Can Adapt for Your Own Code
This is an original prompt for the manual workflow above. Replace the bracketed fields with your project details and attach the relevant code. For a tiny task, remove sections that do not apply.
Help me implement the following task through a structured coding workflow.
TASK
[Describe the required behavior.]
PROJECT CONTEXT
[Language/runtime, framework version, relevant files, existing interface,
and the current behavior or error.]
ACCEPTANCE CRITERIA
[List observable requirements, input/output examples, and edge cases.]
CHANGE BOUNDARIES
[Name the permitted files and any dependency or interface constraints.]
Preserve existing user changes. Keep unrelated cleanup out of this task.
STAGE 1: PLAN
Read the supplied context. Summarize the intended implementation and
identify any unresolved decision that changes externally visible behavior.
Ask a focused question only when that decision blocks a correct result.
STAGE 2: IMPLEMENT
Produce one candidate that follows the agreed requirements and existing
project conventions. Explain any necessary expansion of the stated scope.
STAGE 3: REVIEW
Check the candidate against the original acceptance criteria.
For each actionable finding, provide the triggering input or condition,
expected behavior, and the relevant code location.
Separate defects from optional style suggestions.
STAGE 4: VERIFY
Use [the project's test command or specified checks] if execution is
available and authorized. Otherwise provide the exact checks for me to run.
Distinguish expected results from observed output.
Never report a test as passed unless it was actually executed successfully.
STAGE 5: REPAIR
Address confirmed failures and rerun affected checks when possible.
Limit automatic repairs to two rounds; then report the unresolved blocker.
Do not weaken requirements or tests merely to obtain a pass.
FINAL OUTPUT
Provide the final code or patch, a concise change summary, verification
results, and any unresolved limitation.
Use concise work products; do not invent a dialogue between teammates.
If working as one assistant, do not claim separate agents were launched.If the first implementation misses an important rule, correct the relevant instruction and ask for a focused revision. Restarting the entire process can discard useful evidence about what failed.
Keep Handoffs Small and Specific
For separate chats or automated agents, prepare a handoff containing the current task, accepted decisions, code revision, unresolved findings, and verification status. Include the actual relevant code; a summary cannot substitute for an implementation the reviewer needs to inspect.
For example: “Revision B preserves order and removes blank entries. The Unicode comparison check fails. The function signature must remain unchanged. Review that behavior before suggesting broader changes.” This gives the next stage a defined job.
When a role stops following project conventions, investigate missing project instructions before adding more agents. A rule unavailable in the working context cannot reliably guide the handoff.
I would also name one owner for each writable file in an automated workflow. Reviewers can inspect shared material, but simultaneous incompatible edits create a separate coordination problem. Consolidate the final patch and rerun affected checks after integration.
Does Self-Collaboration Actually Improve Coding Results?
It can be worth trying when a task has several requirements that are easy to overlook, or when an implementation needs a targeted second review. I would use a simpler prompt for a straightforward syntax question or an obvious one-line change.
OpenAI's evaluation guidance recommends letting evaluations determine whether a multi-agent architecture is justified. Additional agents also introduce more handoffs and opportunities for variability.
Compare workflows on representative tasks using the same requirements and verification standards. Record accepted results, regressions, review time, elapsed completion time, and model or API usage where available. Include the time you spend preparing context and repairing generated code.
Historical reporting on AI coding productivity covers METR's early-2025 study, where experienced contributors working in familiar repositories took longer with the tools studied. That result concerns a particular population and tool period, not a verdict on this workflow. In February 2026, METR also cautioned that selection effects made its follow-up data an unreliable estimate of current productivity gains.
For your own decision, compare completed work of acceptable quality. A longer answer, more review comments, or more generated lines does not establish improvement.
Failure Modes to Watch For
Shared mistaken assumptions: several roles can accept the same incorrect interpretation. Keep the original requirements available and resolve disagreements with examples or evidence.
Tests that merely repeat the implementation: if the test author copies a bug into the expected result, a passing test provides little reassurance. Derive expected behavior from the contract.
Review-driven scope growth: an optional rewrite can create more risk than the requested fix. Require a concrete reason for each additional change.
Confusing instructions with access controls: telling a reviewer to avoid edits is an instruction. When the environment supports it, read-only access provides a separate technical restriction. Match execution and write access to the actual task.
Approval without evidence: “looks correct” is a review opinion. It should not replace the agreed tests or a final examination of the patch.
Frequently Asked Questions
Can I try self-collaboration in a normal ChatGPT conversation?
Yes. Request the stages in successive turns and provide the required code and context. Whether tests can run depends on the tools and runtime available in that session; use your development environment when necessary.
Do I need three different AI models?
No. You can assign different responsibilities while using the same model. Different models are another experimental choice, but their additional cost and differing behavior need evaluation rather than an assumption of better results.
Is self-collaboration the same as repeatedly asking for improvements?
A structured workflow defines what each pass must produce and what evidence justifies a revision. An open-ended request to improve code has no equally clear endpoint or acceptance standard.
Can the reviewer replace human code review?
I would use it as assistance. Someone still needs to decide whether the requirements are right, whether the checks are sufficient, and whether the patch belongs in the project. The appropriate review depth depends on the consequences of an error.
Does calling a role a tester mean it executes tests?
No. Inspect the tools used and the execution output. A role label or a predicted result does not demonstrate that the program ran.
Try It on One Bounded Task
Choose a small utility or reproducible bug, write its acceptance criteria, and run one implementation-and-review cycle. Execute the relevant checks and inspect the final change. Keep the extra stages if they catch meaningful defects at an acceptable cost; simplify the workflow if they mainly produce repetition.