Cited

The prompts worth keeping in a file

Most prompt libraries are padding. The dozen we reuse weekly, why they survived, and how to maintain a library that does not rot.

A navy blue fabric gift box with a blank card on top, perfect for special occasions.
Photo: https://kaboompics.com/ / Pexels

Part of AI writing in real work, not in demos

We went looking for our own prompt library recently, the one we built a year ago with good intentions, and found something worth admitting to: it had eighty-one entries and we could name maybe eight from memory. The rest were not bad prompts. We had simply never gone back for them, because going back meant scrolling a long document and deciding whether "Rewrite for executive audience v2" was meaningfully different from "Rewrite for executive audience final". It was faster to retype the instruction from scratch.

That is the failure mode of every prompt library we have seen shared publicly — newsletter lead magnets, GitHub repos with three hundred stars, paid packs sold as productivity tools. They are built to look comprehensive, a different goal from being usable under deadline. A library with five hundred entries is not a tool, in the sense we use the word in AI writing in real work, not in demos. It is a catalogue, and nobody searches a catalogue for the third time in a week. They open a blank chat window and type.

So instead of building a bigger one, we kept a log. For six months, every time one of us reused a prompt rather than writing a fresh instruction, we noted which one. Twelve prompts accounted for nearly all of the reuse — not twelve categories, twelve specific, worded instructions, used again and again on different documents. Here they are in full, along with what they share and the maintenance rule that keeps a library this size from rotting.

The twelve

Cut to length, keep the argument. "Here is a passage of [X] words. Cut it to [Y] words. Keep every claim and every number. Remove only words, transitions and examples — do not summarise, do not paraphrase sentences that are already tight." Used more than anything else on the list, for the reason we go into in why the tools work as editors and fail as drafters: the content is already ours, so a bad result is obvious on the first read.

Neutralise the tone, keep the content. "Rewrite this with the same information and conclusion, at half the emotional temperature. Do not soften the substance, only the delivery. Flag anything removed rather than deleting it silently." For the email written at 11pm that should not go out until it has cooled.

Find the inconsistent terms. "List every case where the same thing is called by more than one name, and every place the same term is used for two different things. Do not fix it, just list it with line references." A pure inventory — the fixing stays manual.

Attack the weakest claim. "Read this document as someone who disagrees with its conclusion and wants to find the easiest place to attack. Name the single weakest claim and say why." Deliberately narrow — "improve this" produces mush, one named weak point does not.

Convert notes into an outline, nothing more. "These are raw meeting notes. Produce a numbered outline of the decisions made and the open questions, no additional commentary, no summary paragraph." Outline only, never the write-up — the write-up is where invented connective tissue creeps in.

Turn a transcript into action items. "From this call transcript, list only the items someone explicitly committed to doing, with who and by when if stated. If neither was stated, write 'unassigned' or 'no date given' rather than guessing." That clause exists because early versions filled in plausible owners nobody had volunteered.

Rewrite a heading as a claim. "Rewrite this heading so it states the conclusion of the section rather than naming its topic." A heading that says something serves a skimming reader better than one that merely labels.

Check a table against its source text. "Compare this table to the paragraph above it. List any number that appears in one and not the other, and any that appear in both but do not match." Proofing, not drafting — the one research-stage task we trust.

Strip a paragraph of hedging. "Rewrite this paragraph removing hedge words — 'somewhat', 'arguably', 'to some extent', 'in many cases' — without changing what is claimed. If removing a hedge would make a claim false, leave it and say why." The exception matters more than the instruction; without it, hedges that were load-bearing get deleted too.

Compress an update for a reader in a hurry. "Rewrite this as three sentences: what happened, what it means for the reader, what happens next. No preamble, no 'I wanted to update you'." Distinct from the compression prompt above — it forces a fixed shape rather than a target length.

Translate jargon for one named reader. "Rewrite this paragraph for [specific person], who does not know [specific term list]. Replace or define those terms; leave everything else close to the original wording." Naming an actual person, not "a general audience," is what makes it usable — a general audience is nobody in particular.

Ask what a reader will quote back. "Read this document as someone preparing to summarise it to their boss in one sentence. Write that sentence. If it misrepresents the document's actual point, say so." A late-stage check that has caught documents making an argument they had drifted away from by the final paragraph.

What the twelve have in common

Every one does a narrow, single job. None says "improve this" or "make this better," because that asks the model to guess what better means, and the guess regresses to an average version of the document type. Every one states its output format — a list, three sentences, a single flagged claim — so the result can be checked at a glance against what was asked for, rather than read closely to work out what happened.

And every one leaves the content to us. Not one asks the model to originate an idea, a claim or a fact. They compress, flag, sort, check and reformat material that already exists — the entire reason these twelve survived six months of real use while the other sixty-nine did not. A prompt that asks for judgement produces a plausible answer you then have to verify line by line, slower than doing the judgement yourself. A prompt that asks for a mechanical transformation produces something you can accept or reject in the time it takes to read it once.

Model drift is quiet, not sudden

Three of the original eighty-one prompts stopped working well at some point in the last year, and none threw an error. The output got worse in a way that took weeks to notice, because each individual result still looked reasonable on its own.

The clearest case was a prompt built around a model's old tendency to over-explain its reasoning before answering. We had written the instruction to work around that habit — "skip the explanation, give only the list" — and when an update made the habit disappear on its own, our workaround started producing oddly truncated output, since we were now suppressing an explanation no longer being offered in the first place. We caught it only because someone happened to compare two runs from different months side by side while cleaning up a folder.

The fix was not clever. It was scheduling. We do not trust ourselves to notice drift by feel, so once a quarter we rerun all twelve prompts against a document whose right answer we already know, and read the output cold.

Where to keep them

The library matters less than where it lives. Ours is a single plain-text file, pinned in the same note app we already have open while writing — not a wiki page, not a shared doc. The rule: if a prompt takes longer than three seconds to find, we retype it instead, and a prompt nobody retypes might as well not exist. The file also stays flat rather than sorted into folders, because categorisation helps you browse and browsing is not what happens under deadline.

The quarterly prune

Once a quarter we go through the file and remove anything not used since the last prune, no matter how good it looked when we wrote it. This is the discipline a five-hundred-entry public library skips, because a library built to be shared has an incentive to look complete rather than to stay short. Ours has the opposite incentive: it has to fit on one screen, because a file you have to scroll is a file you stop opening. Two prompts have been removed this way in the last year, both because the task they were built for stopped coming up, not because the prompts failed — a smaller number than we expected going in, which is itself the case for keeping the library narrow. The twelve that earned their place have kept earning it, month after month, on exactly the documents we described in our notes from testing the writing tools on live documents.

Questions people ask

How many prompts should a working prompt library actually have?
Fewer than you think. Ours settled at twelve after six months of watching which ones we actually reused; everything else we wrote once and never opened again.
Where should I store the prompts I reuse?
Somewhere that opens in under three seconds without a search — a pinned note, a snippet manager, or a single file you keep in the same window as the document you are editing. A folder inside a wiki you have to log into is where prompts go to die.
Do prompts stop working as models change?
Yes, and quietly. A prompt that relied on a model's old habit of over-explaining, or its old refusal to touch certain phrasing, can start producing worse output after an update with no error message telling you why.
Is it worth buying or downloading a large prompt pack?
Only as a source of ideas to steal one from, not as something to use as-is. A five-hundred-entry pack optimised for looking complete, not for being searched under deadline.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.