Calculated Moves
The rule this page runs on — evidence over persuasion — in the artist's own words: I don't beg for grace, I weigh the proof… no stunt, no bluff, just coded proof.
How the work actually gets made. The value is not access to AI — it is disciplined control of it.
In plain language. This page describes how one person makes a large amount of finished work on a phone, with help from AI tools that are never left unsupervised. Every job is split into small, separate steps — one system analyzes, another plans, another checks the plan, another does the work, and another checks the result. No single tool is trusted to do everything, and no result is trusted just because a tool says it worked: the finished file itself is re-checked. Original files are protected before anything is changed, so a mistake can be undone. The rules for all of this are written down, and when something goes wrong the rules are updated so that kind of failure cannot happen again. A person, not a machine, makes every final decision — and the final test for the music is a human ear, not a measurement.
Who this is. An independent artist, machine operator and technician who designs human-directed, multi-AI production systems on a phone. Different models are assigned bounded roles — analyze, plan, audit, execute, verify — and no single AI is trusted to do all five on anything that matters.
What it has produced. A completed 150-track, 11-album catalog, mastered through a measurement-driven remastering pipeline with a frozen, re-certified processing chain; FretNova ∞, a guitar-composition engine built on real theory and playable ergonomics; and an in-progress narrative music video produced clip-by-clip under a written production bible with automated prompt verification.
How it works. Prompts are written as operating specifications: role, mode locks, hard constraints, output schemas, uncertainty rules, and explicit verdicts (PASS / FAIL / INDETERMINATE / APPROVE WITH REQUIRED CHANGES). Originals are protected, results are verified by re-reading the artifact rather than trusting a success log, and the ear outranks the meter.
A controlled pipeline built around specialization, bounded autonomy, and evidence.
The workflow is built on a simple but demanding idea: different models are better at different jobs, and high-stakes changes become safer when those jobs are separated. Each AI gets a defined role rather than one model improvising the entire process from beginning to end.
What gets protected first. Four priorities in a fixed order, and every rule below inherits from them.
Task completion is last. An unfinished task is a minor cost. A destabilized device, a destroyed original, or a false claim shipped to a live site are not.
Twelve steps, numbered from zero. Every step below opens for the full description, and several carry a worked example from project history.
Every project here starts as a conversation, not a specification. Before anything is defined, AI is used the way an unfamiliar field is best approached: what does this domain actually require, what are the industry standards, what is current practice versus what merely used to be, what will need real research later, and which parts of the vision are going to be the most difficult. The output of this step is not a plan. It is an accurate picture of the problem — including an honest read on the hardest parts of it, arrived at before any effort has been sunk into the easy ones.
In practice · A mastering laboratory, a harmonic composition engine, and a 150-track catalog were all specified by someone with no formal training in any of the three. That was possible because the specification came second. The first move, every time, was finding out what the field already knew.
Published figures, prior reports, inherited settings, and anything a previous session concluded all enter as unverified. Before an objective is defined, each load-bearing claim is traced back to the file, measurement, or document that produced it — or it is re-derived from scratch. This is the step that catches the errors nothing downstream can, because every later step is measuring the artifact rather than interrogating the record it inherited.
In practice · An outside review of this page arrived with a computed figure that disagreed with a published one, and recommended correcting the site. The standing rule is that a computed number disagreeing with a published number is first assumed to be measuring something different — and it was: the two counted different categories across different corpus sizes. The published figure stood, unchanged.
The objective, the source material, the standards that apply, what is explicitly out of scope, and what finished will look like — all written before a model is asked for an opinion. Acceptance criteria are fixed at this point so they cannot be quietly relaxed later to match whatever the output turned out to be.
A deterministic baseline before subjective interpretation. The numbers come from tools that return the same answer twice — not from a model's impression of the material. Interpretation is allowed to explain the baseline. It is not allowed to replace it.
A model is asked to find issues, patterns, risks and opportunities in read-only mode. Analysis carries no authority to change anything. Findings are kept separate from recommendations, so a weak recommendation cannot ride in on the strength of a good finding.
A large share of this work is not analysis of the material but investigation of the machinery that will touch it. Documented findings, not assumptions. A processing filter whose interface changed across major versions. An encoder that does not behave predictably as a setting is raised. A limiter that turns out to be measuring a different kind of peak than the one that matters. None of that came from listening to audio, and every one of them changed the approach. Findings are written down with their source, so the next session inherits the finding instead of rediscovering it.
Findings become a specific action plan: what will be done, why each choice was made, the effect expected, and how that effect will be measured. A plan that cannot state its own failure condition is sent back.
The auditing model is given enough independence to disagree, and is asked to look for contradictions, unsupported conclusions, unsafe actions, missing evidence, and failures to follow the specification. Running alongside the technical audit is a provenance check, which asks a different question: not "is this conclusion correct" but "where did this claim come from, and can it be traced." A report can pass every integrity check it runs and still carry a section attributed to a decision that was never made. No measurement catches that. Attribution does.
In practice · Audit is a loop, not a gate. Rejected work returns to Plan, and sometimes all the way to Define — one disputed case was judged under a wrong reading, had its basis rejected, and was redefined a third time before any recount ran.
Sequenced work, not a checkbox. Old backups are cleared first, so the wrong one cannot be restored later. A fresh copy is made of everything the run may touch. The replacement is built and confirmed complete in a staging area, and only then does it displace the original. If a run fails partway, the original is restored rather than repaired — and the live version is never left missing while the problem is worked out.
Only the approved actions, within the authority the task granted and no wider. No installs, no invented processing, no destructive shortcuts, no quietly extending the job because something else looked worth fixing. Resource limits are part of the contract rather than an afterthought, because the machine has to survive the run. The agent reports what it did against what it was approved to do, and the difference between those two is itself a finding.
A clean exit proves nothing. Verification means evidence drawn from the artifact itself: reading the result back and comparing it against what was supposed to be produced, rather than trusting a report that says it worked. Completion claims are treated as unverified until the artifact is re-examined. This standard exists because a batch job once reported embedded artwork it had not written.
Release, revise, or discard — on evidence and human judgment. Human correction is not confined to this step; it enters wherever it is needed, and in practice it enters mid-run. Work is handed off and left to run, so the rules for interrupting it are two-tier: halt and ask when the ambiguity is consequential, correct and continue when the defect is trivial. The test for the second tier is written down, and if there is any real chance the human would have chosen differently, it halts. Definitions get overruled while the run is still open: a wrong reading is redefined and re-run, not patched at the end. Every session produces a report regardless of outcome, most serious findings first, because burying a serious finding at the end is itself a defect.
This model appears across the technical and creative work. In audio, an analyzer does not automatically become the mastering executor. In software, a coding agent does not receive unrestricted permission to reorganize the workspace. In content strategy, AI recommendations are tested against platform behavior, audience fit, catalog goals, and artistic identity rather than accepted because they sound persuasive.
The steps are drawn in order because they execute in order, but the run does not end where it stops. Audit can return work to Plan. Verify can return it to Research. A failure does not just get fixed — it amends the specification that permitted it. A near-miss that could have taken the device down produced a root-cause diagnosis, which produced two revisions of the operating rules, which now close that class of failure permanently.
Two rule changes are on the record. A session limit was loosened after fourteen clean runs showed it was costing more in restart overhead than it was saving. And when a genre definition was corrected mid-analysis, the first batch was not patched — it was deleted and re-judged in full, which moved the tag from 3 of 14 songs to 9 of 14. That jump is the rule change, not a drift in judgment, and it is recorded as such.
An executing agent gets the authority its task requires and no more. It does not choose its own scope, does not extend the job it was given, does not fix things it was not asked about, and does not act on inference. Anything outside the stated scope is raised rather than performed.
The same principle covers the machine. The whole system runs on a phone, so correctness is only half the specification; the other half is what the hardware can physically survive. Limits are stated before the run starts rather than discovered during it — a floor below which the work stops, thresholds that end a run rather than push the device toward failure, and a hard prohibition on installing anything. On this hardware an install can silently fall back to compiling from source, which costs hours and fills memory that is never released. That has already happened once. A missing capability is a finding to report, not an obstacle to route around.
The agent is capped so that it fails before the device does. A crashed agent is recoverable. A locked phone is not.
Prompts are written as operating specifications, not casual requests.
The strongest AI skill here is the ability to convert a complicated goal into a bounded, inspectable contract. Many of these prompts function more like technical specifications or standard operating procedures than ordinary chat instructions. They tell the model what success means, what evidence is required, what actions are forbidden, how uncertainty must be reported, and what files or records must exist when the work is complete.
A recurring design pattern is the use of explicit verdicts. Instead of asking an AI whether something "looks good," the prompt may require one of a limited set of decisions: PASS, FAIL, INDETERMINATE, NOT OPTIMAL, APPROVE, APPROVE WITH REQUIRED CHANGES, or REJECT. Each verdict must be supported by traceable findings. This reduces vague approval language and makes disagreements between models visible.
Prompt engineering is one part of this, and on its own it is the least interesting part. The larger practice is systems design, context engineering, model selection, orchestration, research, specification design, evidence management, adversarial review, controlled execution, verification, and the retention of human decision authority. A well-written prompt inside a badly governed system still produces work nobody can check.
A long conversation is not a free resource. Every session accumulates superseded instructions, abandoned approaches, and stale state, and a model reasoning across that residue will eventually act on something that is no longer true. Context is therefore managed deliberately, as a threat to project integrity rather than as an inconvenience.
Context bloat is not a performance problem. It is an accuracy problem, and it is handled like one.
One human director, several reasoning systems, and a fixed rule about who is allowed to decide.
The pipeline in section 01 is the production mode. It is not the whole system. Around it runs a wider layer of work: research, source discovery, independent analysis, technical and strategic advice, coding, troubleshooting, document consolidation, specification alignment, contradiction detection, fact-checking, adversarial review, alternative approaches, and independent verification.
The working stack. Claude and Claude Code · Codex · ChatGPT · Grok · Perplexity · Gemini · DeepSeek. Claude often serves as the project workspace for long-running work, with Claude Code and Codex as the agents that execute, but they are not the entire system — other models may work before, beside, or after them. Not every model is used on every task, and none of them holds a permanent job title; what each is asked to do changes with the problem. Naming them describes a working method, not an endorsement, partnership, or affiliation in either direction.
HUMAN DIRECTOR
↓ ↑
MULTI-MODEL LAYER
Claude · Claude Code · Codex · ChatGPT · Grok
Perplexity · Gemini · DeepSeek
research · critique · troubleshooting · fact-checking
· competing solutions ·
↓ ↑
CONTROLLED EXECUTION PIPELINE
Orient → Verify inputs → Define → Measure → Analyze →
Research → Plan → Audit → Back up → Execute → Verify
↓
HUMAN ACCEPTANCE
The objective is not to collect the largest number of AI answers. It is to reduce dependence on any single model's assumptions, and to improve the quality of the decision a human then has to make.
FretNova is the clearest example. Some development problems on it went past what the primary workflow could reliably resolve — an architectural question that kept producing plausible but wrong answers, an implementation that failed for reasons none of the obvious explanations covered. At that point the problem was widened rather than repeated: several systems examined the same failure independently, their reasoning was compared and challenged against each other, the parts that survived were consolidated, and the result was carried back into the main project. Some of those problems took sustained work across several systems over days. None of them were solved by a better single prompt.
For consequential work, independent agents may analyze the same source without seeing one another's findings. Their results are consolidated separately, the consolidation is independently verified, and execution occurs only after that review chain is complete. Roles remain deliberately separated so that no stage silently expands its own authority or certifies its own work.
Nothing in that chain lets a stage grade its own work. The analyst does not consolidate. The consolidator does not verify. The verifier does not execute. The executor does not certify. And the documents the chain runs on were themselves written only after the source information had been independently checked and reconciled across the wider stack, in a deliberate order, with the human's design decisions, rules and constraints settled and written in before any agent was given a job to do.
The layers are not interchangeable, and that is the point. Planning, execution and record-keeping are held apart: whatever decides what should happen does not carry it out, and whatever carries it out holds no authority to decide what should happen. Execution is bounded by a written specification rather than by judgment in the moment. A third responsibility sits alongside both: when something structural changes, the specification is updated in the same pass, unprompted, and the new version says what it supersedes.
None of the three can quietly take on another's job. The planner cannot execute, because it has no hands. The executor cannot expand its own scope, because the scope arrives in a document it must read first. And neither can amend the rules, because a proposal is not active until it is approved.
None of this is choosing between AI-generated options. The human defines the problem, sets the objective, supplies or creates the source material, establishes the constraints, designs the workflow, decides which systems are involved, weighs their disagreement, directs the revisions, determines what evidence is sufficient, accepts or rejects the result, and maintains the authoritative record everything else is measured against.
On the music side that matters more, not less. Songs are not limited to text-prompt origins. They may begin as lyrics, guitar playing, arrangements, compositional ideas, or material written years before any of this existed. AI enters afterward, in transformation, development, analysis, refinement, production, or verification, depending on what the work actually needs.
It does not eliminate AI error, and nothing here should be read as claiming it does. Separate commercial systems are not independent in any scientific sense; they share training data, methods, and blind spots, and they can be wrong together. What the method offers is narrower and more defensible: reduced dependence on a single model, preserved evidence, constrained execution, and a human who stays accountable for the result. Corroboration between systems is a reason to look closer, not a proof.
The workflow assumes that every model can be wrong, persuasive, or incomplete.
Multiple AIs are used not to create an illusion of consensus, but to expose disagreement. A second model is most valuable when it is given enough independence to challenge the first. In critical workflows, the reviewer is instructed to look for contradictions, unsupported conclusions, unsafe actions, missing evidence, and failures to follow the specification.
A claim of success requires evidence from an actual check, not from the fact that a command finished without complaining. Every kind of claim has a kind of proof attached to it before the work starts, and the proof has to come from the artifact rather than from the report about the artifact.
That is a harder standard than it sounds, because most checks verify something adjacent to what was actually asked. A syntax check does not prove a page renders. A copy that succeeds does not prove the copy is identical. An archive that builds does not prove it holds the right files. So the instruction is procedural rather than technical: ask what could still be broken if this check passed — then go and check that instead.
Some judgments cannot be settled by a tool — whether a song belongs to a genre, whether a recurring image is a motif. Those are governed by a written rubric rather than by taste, so that the same song judged months apart gets the same answer.
Prompts require evidence, explicit uncertainty, exact file references, and structured outputs. Unsupported completion claims are treated as defects.
Checklists, indexes, manifests, current-state handoffs, and clean workspaces prevent an agent from silently changing the scope or acting on stale files.
What the system produced when it was run end to end.
The catalog is 150 tracks across 11 albums, released under the artist's own label. Every track was measured, assigned a treatment, processed through the frozen chain, checked against both a WAV ceiling and a codec simulation, tagged, and reconciled against a release manifest. One constraint held throughout the mastering pass: all 150 clear the gates before any of them ships. No shipping the easy tracks while the difficult ones stayed unresolved.
| Workstream | Result |
|---|---|
| Mastering | All 150 tracks clear the codec true-peak gate at ≤ −0.3 dBTP. The 32-track problem set was resolved by re-engineering and re-certifying the final stage of the chain rather than by lowering the ceiling. |
| Metadata | Full ID3v2 tagging across the catalog: artist, album, label, title, track number, year, composer, credit, artist URL, embedded lyrics, and cover art. |
| Lyrics | 150 lyric masters cleaned, reformatted, disambiguated, and re-embedded, each verified against the audio it belongs to by hash, with platform-specific derivation sets built for the distributor and the lyric provider. |
| Artwork | Embedded art verified by extracting the image back out of every file and comparing hashes, after a batch reported a result it had not achieved. |
| Sequencing | An album was resequenced after release planning changed: track numbers rewritten and filenames prefixed, because the target player ignores track tags in folder view and orders by filename. |
| Release ops | A single track was submitted as a pilot for the distributor's audio-swap process and confirmed before any batch swap was authorized. A storefront kit of per-album descriptions, per-track notes, tags, and upload checklists was generated for the direct-sales channel. |
What the mastering gate measures, in numbers. The gate this lab publishes is the codec true-peak ceiling: every one of the 150 tracks clears ≤ −0.3 dBTP measured after codec simulation, not on the WAV alone — the encoder is where true-peak overshoot actually appears, so that is where it is tested. Elsewhere on this site, references to professional loudness and true-peak standards point at this measurement. The integrated-loudness targets and the per-track reports behind them are documented but not yet published; they are available on request.
Two decisions are worth naming because measurement alone would not have produced them. A set of 108 songs ships with no structure labels in the lyric-provider format; that was verified as correct rather than fixed, because the source lyrics never contained labels and inventing them would have been fabrication. And a closing track was moved to a different album once it was identified as sharing verbatim lines with an earlier song — a repetition that reads as deliberate catalog architecture in one position and as an error in another.
The catalog analysis is the better demonstration of what this page claims, because nobody watches it happen. The work is defined once, in writing, together with the record of how far it has got. A session starts, determines its own position in that record from the document rather than from a person, does the next piece of work, writes down what it did, and stops. No stage is named by hand. Sessions have a budget and end well before they reach it, and an agent that notices it has contradicted itself or lost track of what it already read is required to stop on that basis alone.
Every session produces a report, including a session that changed nothing. A system that demands a report from a run with nothing to report is a system that expects to be audited.
Christopher Jager is an independent artist, producer, systems designer, and advanced AI practitioner. He builds human-led, multi-AI workflows for music, software, audio engineering, visual production, and platform strategy. His systems separate analysis, planning, execution, audit, and verification; protect original files and artistic intent; and use evidence, structured prompts, and independent review to make AI output dependable rather than merely impressive.
Human-led music and technology, built through advanced prompting, multi-AI orchestration, and independent verification.