We put all of our engineering rules in one place, gave each one an identifier, and made every AI assistant on the team, and the reviewer in CI, read the same text. Such a simple change resulted in many positive outcomes, some unexpected.

Executive summary

If you lead an engineering organisation, you have probably noticed that AI coding assistants are only as good as what they are told. Ours were being told different things. Every repository had its own instruction file, written by whoever touched it last. The files drifted apart, and away from the code. Behind this sits a new part of every developer’s job: context manager. The work now starts before the first line of code, by making sure the right rules and facts make it into the session. Doing that well comes down to three habits. Keep team rules and system facts in files the tools load on their own, not in prompts that stay private to one person. Load only what applies to the code being touched, so the rules that matter are not buried. And when something had to be typed into a prompt to get the right result, move it into a file once, for everyone.

So we did something simple. We put every rule in one repository and gave each rule a short, stable identifier. Repositories now carry only a pointer to the central rule repository. When a developer starts a session, their assistant fetches the current rules: the general ones, the ones for that repository, and a few that load only when matching code is touched. The reviewer in CI loads the same set, and every finding it posts names the rule it rests on. Changing a rule is a pull request, reviewed by the people it affects. Facts about a system stay in that system’s documentation; the rule repository says only what to do.

A few things happened. The one we hoped for: a rule we agreed in April now applies to everyone’s pull request in October, and nobody has to remember it. The question “why does this PR still not follow rule X?” has stopped coming up. This results in fewer human review rounds. We also managed to reduce token usage – only what’s relevant and is not a well known fact is carried in the session. The one we did not plan: developers started proposing rules in the first week, because they saw the benefit: a well-crafted rule is added in one place, and a rule the team agrees on in a retro becomes a guardrail the assistants and the reviewer enforce from then on.

We can also measure it. Nine repositories were switched over in two days. The reviewer’s audit mode told us which rules each review applied, so we knew the right ones were loading before we trusted it. It also cut token spend. Each review loads only the rules that apply to it: 13 to 60 of about 300, or 4 to 20 per cent of the set. Assuming rules of similar length, that is 80 to 96 per cent fewer rule tokens than loading the whole set into every review. The rule context settles at about thirty thousand tokens, loaded once, and a review costs between a third of a dollar and two dollars, rising with the size of the change, not with the number of rules. Coverage improved at the same time: in about 80 reviews none failed to load a rule file, so every review carried every rule that applied. Pruning cut the set further, because conventions that restated well-known public guides became one-line pointers to those guides.

For a regulated business, the point is not that the assistant is clever. It is that its behaviour is now governed: rules are versioned, reviewed, attributable and testable, and a review comment traces back to a decision the team made. We also wrote down what the mechanism does not do, so nobody assumes more than it delivers. The rest of this article is for engineers who want to adopt it, challenge it or improve it.

The developer’s new job: context manager

Here is a useful way to think about working with an AI assistant: it produces what is likely, given what is in its context. Whether that matches what your team wants depends entirely on what was in the context when the code was written. And that depends on the developers’ understanding of the business domain. So every developer has quietly picked up a new responsibility. Before the assistant acts, all the facts that matter have to be in front of it: the repository’s conventions, the risks the team has accepted and does not want flagged again, the trap that cost a week in March, the detail that this environment runs a single database instance so the reader endpoint is really the writer.

A developer who knows all of this still has to get it into the session, and the obvious way is to type it into the prompt. That does not hold up. It is tedious, so people stop. It is easy to forget, and the fact you forget is the one that mattered that day. And it is private: what a senior engineer types into their prompt never reaches the junior engineer’s prompt, or the reviewer’s.

So we moved the context out of prompts and into files that the tools read on their own. Rules live in the central repository and load when a session starts. Facts live in each repository’s documentation, and the rules point at them. Narrow rules are tied to file paths and load only when a matching file is read, so a session editing Terraform or Terragrunt does not carry the Go rules around. The procedures that used to be tribal knowledge, like how to run the CI checks locally before pushing or how to turn a review finding into a rule change, became skills: short instruction files the assistant loads by name.

What is left of the context-manager job is the good part: noticing when something is missing from the files and adding it there, once, for everyone. The assistant’s behaviour stops depending on who is typing, which is what makes it governable. And the files become the place where the team’s knowledge lives, with history, review and an owner.

Where we started

Like most teams, we began with an instruction file in every repository. Each one had the same shape: a block of general rules (the security baseline, how we write code, how we handle branches and pull requests) copied from the last repository, followed by the rules specific to this one. Above them, a workspace-level file held the estate-wide rules and a map of the repositories, and the CI reviewer had its own prompt with its own checklist.

It worked for a while, and then it did what copies do. The general block in one repository was three edits behind the one next door. A rule we tightened in the backend never reached the infrastructure repositories. Facts crept into the rule files: ports, commands, directory maps, right on the day they were written and wrong a month later. Nobody owned the whole set, so nobody pruned it, and the files grew. The assistant read all of it at every session start, relevant or not, and the reviewer read something else.

How it evolved

The change came in three steps, and each one was small enough to ship on its own.

One repository, pointers everywhere else. We created a central repository for the rules and reduced every other repository to a pointer: a one-line instruction file that imports the rule index, the hook configuration that fetches the rules, and two ignore entries. The per-repository rule files were deleted in the same pull requests. The reviewer in CI gained a flag that makes it check the central repository out and load the same index a developer’s session loads.

Deduplicate and reorganise. With everything in one place, the duplicates were obvious and went first. Then we sorted what remained into three tiers: general rules for every repository; one file per repository for its conventions, accepted risks and traps; and path-scoped rule files that apply only to matching code. Facts moved out to the repositories’ documentation, and the rules now point at them. Every rule became one bullet with a stable identifier at its start, so a review comment can cite it and a search can find every place it is used. A rewritten rule keeps its identifier, a retired one is never reused, and a small script in CI enforces both.

Prune to the pointer. A good share of the rules restated public standards the model already knows: a well-known Go book, the PostgreSQL “don’t do this” page, the usual security checklists. We replaced those paragraphs with one line each: which standard we follow, and that findings should name its item. The model knows the alternatives; it only needed to be told which one we chose. The rule set got shorter and the reviews got sharper.

Alongside the rules, the procedures that used to be tribal became skills: how to run the CI checks locally before pushing, and how to turn “from now on, when working on X, always do Y”, said in a session, into a rule change in the right tier.

Loading the right rules at the right time, every time

Having one source is half the job. The other half is making sure a session gets exactly the rules it needs, no more and no less, and that it really gets them.

The same rule text reaches the developer’s session at start, the CI reviewer at review time and the human reviewer in the thread; rule proposals and audit metrics flow back into the one repository.

At session start. A hook runs when a session opens in a repository. It fetches the central repository into an ignored folder as a shallow, read-only cache, resets it to the main branch so no local edit survives, writes the index of rule files that apply to this repository, links the path-scoped files where the assistant looks for them, and prints one status line saying what loaded. If anything fails, the line says so loudly and edits are blocked until it is fixed. Edits inside the cache are blocked too.

Only what applies. The general rules load for every repository. The repository’s own file loads only there. Path-scoped files load lazily, when a matching file is actually read, so a session editing Terraform never carries the Go logging rules and a session in the UI never sees the database rules. The rule context a review carries is about thirty thousand tokens, loaded once, and it does not grow with the size of the change.

Across repositories. Developers work in a workspace folder that holds the repositories side by side. The workspace has its own rendered rule file, so a session started inside any repository reads the general rules as an ancestor. A guard watches edits: an edit in a repository whose rules have not been read in this session is blocked, and the guard names the files to read. A change that spans the backend and the infrastructure repository loads both rule sets, in the order the work touches them.

In CI. The reviewer is a reusable workflow every repository calls. With the flag on, it checks the central repository out, loads the same index, and reviews the diff against it. Each finding carries a severity, a category and the rule identifier it rests on, then the problem, why it matters and a concrete fix. A finding no rule covers omits the identifier rather than inventing one.

Proving it loaded. An audit flag makes the review summary list every rule applied and every rule file it could not read, and the metrics of each review are kept as a build artifact. During the rollout we read those artifacts per repository before trusting the setup. We keep the flag on in our busiest repository as a canary: if a tooling update ever stops the rules from loading, that is where we will see it first.

The way of working: a builder and a devil’s advocate

The rules are half of it. The other half is how we set two differently prompted assistants against each other.

The developer’s assistant is a builder. It has the developer’s intent, the decisions made in the conversation, and the repository’s rules, and it is asked to make the change work. Like any builder, it is fond of its own work: it will defend a design it has just produced, and it reads its own diff assuming the diff is right.

The CI reviewer is prompted as the opposite. It reports only problems, never praise. Its bar is anything that could cause incorrect behaviour, a security exposure, data loss, a test failure or a misleading result, including things it is not sure about, which it has to mark with a confidence level. It is not allowed the cheap findings: it may not ask for a comment to be added, it may not repeat what an earlier round already raised, and it may not wander outside the diff except to check a specific hypothesis with one or two lookups. It skips vendored and generated code. And it reads the same rule text the builder had, so when the two disagree, the disagreement is about the code, not about the standard.

In practice the loop looks like this. The developer and the builder make the change and run the repository’s checks locally, with the rules in context. The pull request opens; the devil’s advocate reviews the whole diff and posts its findings, each anchored to a line and a rule. The developer and the builder answer every finding in its thread: a fix, one commit per round, or an explicit accepted-risk answer that names the ticket, never a silent resolve. The human reviewers come last, and they see a diff that has already survived an adversary, with the mechanical findings gone and the design questions left.

Two things make this work. It works because the roles are asymmetric: the builder carries context the reviewer does not have, and the reviewer holds a stance the builder cannot hold about its own output. It works because both are bound by the same text, so the review checks the team’s rules and not the reviewer’s taste. But be careful while the design is still moving: keep the pull request as a draft the reviewer ignores, write the scope note and tests first, and open it for review only when the adversary has a finished thing to attack.

One more thing the rule book gives the reviewer is the rest of the change. Many changes require PR in more than one repository, for example the application half lands in one PR, the infrastructure or configuration half in another. A rule makes every such PR link its siblings, and CI searches the organisation for open PRs carrying the same ticket code and hands the list to the reviewer. The reviewer reads each sibling’s description and, when the diff it is judging depends on something there, that sibling’s diff too: the secret the code now expects, the permission the role must have, the config key the manifest must set. So instead of a conversation that starts with “did we also change the infra side?”, the review either confirms the dependency is delivered or reports exactly what is missing. This greatly reduces the number of review/respond rounds.

Rules by reference: point at a known source instead of restating it

Not every rule has to be written out. When a well-known public source already says it, one line naming the source and the team’s position is enough: the model knows the source, it only needs to know which one you chose and where you differ. The trick is to forbid the bot from following the link, since when it’s well known, it’s part of the model’s training.

A reference works when the source is public, stable and widely known, so the model has seen it many times: a recognised language style guide, an OWASP cheat sheet, the PostgreSQL wiki’s “Don’t Do This” page. It also needs the team to follow the source as written, or with a few named exceptions, and the source to have item names a finding can cite.

For example:

- **DB-PG-01** **PostgreSQL anti-patterns.** Schema and query code follows the PostgreSQL wiki page "Don't Do This" (https://wiki.postgresql.org/wiki/Don't_Do_This). Exception: `timestamp` without time zone is allowed only in the `reporting` schema, which stores local business dates. Findings name the page's item, e.g. "[DB-PG-01] Don't use money". Apply the page from training knowledge; never fetch the URL.

In our rule set, lines like this replaced paragraphs that restated the same guides less precisely than the guides themselves. UI reviews are the clearest case: instead of passing about 10,000 lines of React best practices into every review, we point at the guide in a rule of a few lines, which cuts the tokens spent on that guidance by more than 99 per cent. Before trusting one, test it: with the rule not loaded, ask the assistant to list the source’s main items. If it cannot, the source is not as well known as you think, and the rule should be written out.

Adopting it in your team

If you want to try this, here are the decisions worth making before you write anything.

  1. One repository or two. If you already have a repository of shared CI workflows, put the rules and hooks there with path-scoped code ownership. If you have nothing central, create one repository for both from day one. We ended up with two and would merge them given a clean start. A layout that works, simplified from ours
  2. The identifier scheme. A scope prefix, a topic, a number: G-SEC-05, TG-02, UI-NEXT-03. Identifiers are never reused; a check in CI enforces uniqueness and retirement.
  3. Rules versus facts. A rule says what to do and points at the documentation; a fact (a port, a command, a layout, a topology) lives in the repository that owns it. Mixing them is what makes rule files rot.
  4. The tiers. A personal preference stays on the developer’s machine. A team rule is a pull request to the central repository. A fact goes to documentation. Decide this once and make the assistant ask which tier applies whenever someone says “from now on”.
  5. The threat model. Write down that the guard is for mistakes, not malice, and list the accepted gaps, before the first review round rather than during the fortieth.

Then the order of work, which took us about two weeks of elapsed time with one person on it part time:

  1. Build the central repository on a draft pull request: the rule files with identifiers, the session-start hook that syncs the cache and writes the index, the guard, the workspace initialiser, and a test suite that runs the hooks against temporary repositories. Live-test it on your own machine before any review.
  2. Give the CI reviewer an opt-in flag that checks the rule repository out and loads the same index, and an audit flag that lists the rules applied. Wire the reviewer’s own repository first.
  3. Pick one repository and one developer as the pilot. The developer runs the onboarding steps cold and reports what surprised them; fix that before going wider.
  4. Roll out with one small pull request per repository, all on the same day: pointer file, hook configuration, ignore entries, the reviewer flags, and the deletion of the per-repository rule files that are now central. Humans merge; the assistant never does.
  5. Leave audit on until each repository has a few clean reviews, read the metrics, then turn it off everywhere except the canary.
  6. Schedule the first pruning review for a quarter later, with the applied-versus-cited data in hand.

The single most useful habit afterwards is the one from the context-manager section: when something had to be typed into a prompt to get the right result, it belongs in a file, once, for everyone.

Leave a Reply