Prompt Library
Conversation Project Kit: Turning One Long AI Chat Into a Source-Cited Archive

Part 6 of the thread Prompts for plans, apps and side hustles
- One prompt plus a folder template turns a long AI conversation into a numbered archive where every claim traces back to the exact turn it came from.
- It works in two phases: paste everything first, analyze only after you say
TRANSCRIPT COMPLETE. - Turns get
T-####IDs and datapoints getDP-####IDs. They're never renumbered, so citations don't break. - Python 3.9+, standard library only. The download is free.
Some AI conversations are worth more than the chat window they live in. I had one that turned into a book project, and I quickly figured out that scrolling back through a 75-minute chat to find "where did that idea come from?" doesn't scale. So I built a kit that turns one long conversation into a folder: numbered, source-cited, and readable by a future Claude or ChatGPT session that has no memory of today.
It's a prompt plus a small Python toolkit. The project it was built on is the one in Measuring the Flourish. That archive ended up with 315 datapoints across 82 turns, each one citing where it came from.
Why not just summarize it?
Because summaries lose provenance, and a book needs it. Six months from now, the only question that matters about any sentence is where it came from.
The design notes are blunt about the other trap: extracting while you're still pasting. It sounds efficient, and it's the most destructive thing an assistant can do here. A claim in turn 12 often gets reversed by turn 180. Themes invented from the first 10% of a conversation will be wrong. And analysis eats the same context window the paste needs.
So the kit works in two phases. While you paste, the assistant banks each chunk to disk word for word and replies with three lines of counts. That's it. No commentary, no "interesting point here!" Only when you send TRANSCRIPT COMPLETE does it start thinking.
The folder layout
Every project is a copy of the template:
| Folder | What's in it |
|---|---|
1. Transcript |
The conversation, split into turn-numbered parts |
2. Data |
datapoints.jsonl, the single source of truth, plus entities and the taxonomy |
3. Datapoints |
Readable views generated from the JSONL: a quick index, by theme, by type |
4. Synthesis |
Summary, threads, contradictions, open questions, book outline |
5. Scripts |
The toolkit |
6. Assets, 7. Meta |
Supporting files, logs, changelog, the assistant's own notes |
97. Raw Source Captures |
Your pastes, untouched |
98. Versioning |
Hashed snapshots, one per chunk |
99. Archive |
Superseded files, since nothing gets deleted |
The numbers make the folders sort in workflow order. The 97-99 folders are reserved and mean the same thing in every project, so anyone landing in an unfamiliar archive knows where the evidence and the history live. Every file stays under 1,450 lines, enforced by a lint script, because past a certain size a model reads the file but stops using the middle of it.
Stable IDs
This is the part that makes it trustworthy rather than tidy. Turns get T-#### IDs and datapoints get DP-####. They're append-only: never renumber, gaps are fine. Renumbering would break every citation ever written against the old numbers.
Each datapoint is one self-contained sentence that cites at least one real turn. If it includes a quote, the validator checks that the quote is a character-exact substring of the turn it claims. Assistant turns get tagged ai-generated, so at draft time I can tell my words from the model's.
The Markdown views are generated from the JSONL and marked as not for editing. Edit one by hand and your change vanishes on the next build, which is the right failure mode because it's loud.
Starting one
Unzip the kit into the folder where you want your archives to live. Open a Claude session with that folder connected and tell it to read _CONVERSATION-ARCHIVE-PROMPT.md and adopt the contract between the BEGIN and END markers. It asks what the conversation is about and when it happened, proposes a descriptive folder name, and then runs:
python "conversation-project-kit/new_project.py" --name "Attention Scarcity and the Bandwidth Analogy" --about "Whether attention behaves like an economic resource." --date 2026-08-21
That creates the new project in the current folder, or wherever --dest points. Then you paste in chunks. If the session fills up, open a fresh one, give it the prompt again, and resume.py prints the whole project state in about thirty lines. The folder is the memory, not the chat.
Tip
In the ChatGPT web UI, select the conversation on the page and copy that. Don't use the per-message copy buttons. They drop the
You said:/ChatGPT said:labels, and without them the transcript can't be split into turns.
Known limits
- Speaker labels are load-bearing. No labels, no automatic turns. The ingest script catches this on the first chunk, but the fix is manual.
- Seam detection is a heuristic. Overlap and gap warnings are advisory. A human still needs to eyeball the first and last words of each chunk at the end.
- It can't judge quality. The validator proves a quote is exact and a turn exists. It can't prove a datapoint is worth keeping.
- The YAML fallback is basic. Without PyYAML, the taxonomy loader uses a regex that handles the template's shape and not much else.
No embeddings or vector database, on purpose. For a handful of conversations, grep over a quick index is faster, deterministic, and still works unchanged in three years.
Download the Conversation Project Kit. It has the prompt, the README, the design notes, new_project.py and the full template.