İçeriğe geç / Skip to content / Zum Inhalt
Ahmet Balaman LogoAhmet Balaman

Consistency in AI Design: The Critic Loop

Ahmet Balaman
Claude DesignDesign SystemArtificial Intelligence DesignPrompt EngineeringDesign LoopClaudeVibe CodingToken CostSub-AgentAI Slop

It's not hard to make a beautiful design for artificial intelligence. It's hard to produce the same beauty twice. You hit it once, you give it the same request the next day, there's something else coming. I'm telling you where the discrepancy comes from and the two concrete solutions: the independent critic cycle and the design that you like.

Where does the discrepancy come from?

The root reason is in one sentence: "The model notes its own homework. " He creates a design, and then he says, "Okay, I did it." In fact, he may have skipped half the bridge, but he's still going to evaluate it.

You can test it yourself, you can show a design that it produces and ask, "What's wrong with this?" Then ask the same question four times, and you'll get five different answers, because it's not a constant measurement, and it's stuck in something else every time.

If you can't use it for a customer, a series of content, or a brand, that design is an accident.

How does the critique cycle work?

The solution is to take the evaluation from the manufacturer. The method is: you give the model a reference and a measurement, it produces model design, and then a couple of independent critics with the "tested context** come in and test the design mercilessly. If it is found missing, the work returns and the cycle begins again.

The "fresh context" part is important. Critics do not see the production process; they only see the outcome and the measurement. A reviewer who sees the process tends to defend his own decisions.

Three roles in practice work well:

  • Brief critic: "Did you really do what you were asked to do?" compares the articles one by one.
  • "System critic: "Does this print fit the design system?" Palet, typography, space rules.
  • "Working critic:** The alignment, the interval, the ratio, the receded image itself.

The result is that pleasure becomes a checklist, which sounds cold, but that's exactly what works: you can't recreate something that can't be measured.

To give you an idea of what a real tour looks like: three failures in the first round (no logo, no color, one color is broken down), the brand sign is too small in the second round, the second accent problem in the third round... a job can go up to ten tours. If you think about why three separate eyes look at each round, the output is so different.

Turning the design you like into a rule

The second solution is more permanent. Instead of giving a design that you like to look like this, you tear it up into pieces and make it a rule.

Here's what's going on.

  1. Find the design you like.
  2. "Do not describe it. " Not "Ferah stops" but "three times the body writing."
  3. Discuss which part is really important, not every detail has to be a rule.
  4. Write the rules in a file.
  5. "Scrip tests that can't be passed by a faulty copy. "
  6. Re-create the original just by looking at your rules.

Step six is the heart of the job, and everything that doesn't keep up when you rebuild is a rule that you forget to write, so you add it and try again, and in a couple of rounds, your rule file becomes really complete.

There should be concrete proportions in the rules: measurements such as gold rates (1.618), 60-30-10 rules for dispersion of color, shade and corner radius values, named for. "Let modern appearance" is not a rule.

When this file is matured, you can turn it into a reusable ability and call it every time. A well-prepared talent asks you what it lacks before it starts producing: who is the subject, what is the background, what is the background.

How do you write a critic?

The criticism works depending on what you ask him. A critic asks, "Is this design good?" gives you an unusable answer. There are four parts of the working critic:

  • "Enter:** Just results and criteria. Don't give away the production process, the tried variants and the model's own reasons.
  • A critic of both typography and hierarchy makes them both superficial.
  • ** The measurement:** has to be a duo. "Could be better" is not a result; "passed" or "stayed" is a result.
  • ** Stay format:** which rule, what element, what measure was broken.

Let me put two outputs side by side to see the difference:

  • Weak: "The title might be a little big, so it's good to check the balance."
  • Useable: "KALDI ا typography scale. Main title / body ratio 3.4, rule 2.5: page title."

The second is a direct, practical directive; the first is a new argument. When you force your critics to write in the second form, the number of cycles drops because each round closes with a concrete correction.

And the rule is, let's not let the critic suggest correction, just diagnose.** The manufacturer should make correction. When you put the diagnosis in the same place, the critic starts defending his solution and loses independence.

What should be found in the rule file?

Your rule file should be made up of concrete values.

  • Tipography scale: title and body punch values, ratio between them, line height, letter range.
  • ** Space scale:** Single base value and its multiples. A 4 or 8-based scale alone prevents most of the mess.
  • The 60-30-10 approach is a good start.
  • The ladder:** how many layers are there, what is the shadow value of each layer?
  • ** The corner radius:** is used by a number of different values. If it's more than two, it has to have a motive.
  • "Move:** In and out times, used softening curves.

The measure of one rule is: can two people read the same rule and come to the same conclusion?

How do you write a test where the wrong copy can't pass?

This is the most skipped step in the process, but the most decisive step, the test is a complete control that captures the violation of the rule.

Three types of tests practically work:

  1. "The height of the text must be between 1.5 and 1.7." If it's not worth it, it stays.
  2. "Set test: "No more than two different corner radius should be used on page."
  3. "Existence test:** "There must be a full call for action in each department."

The way to know if the test is good is: "Prepare a broken version and see if the test catches him. " If it doesn't catch him, it's not a test, it's a wish.

How much does this cycle cost?

There's a price to be honest with you, because there's a multistep review cycle that can spend a million tokens on a single job, a job in shared samples, two or three million token bands.

The way to reduce the cost is to divide the business: to turn critics over to small and cheap models and only keep the actual production in a strong model. In one example, most of the load is covered by the smaller model, most of the load is covered by the middle model. This distinction is taking the cost seriously down.

My suggestion is clear: "Don't use this cycle in every business.** Waste for a one-time visual. The place it's worth, the things you're going to use over and over again: templates, carrier pages, serial content formats, brand assets. Once expensive, then you can use it for free a hundred times.

When shouldn't he use it?

  • If you don't know what you want yet, the cycle does not clarify an uncertain measure; it makes uncertainty more expensive.
  • If the work itself is in the discovery phase, try a few cheap and fast directions first, open the loop after the direction is clear.
  • If the measure is purely subjective, "I like it" is not something a critic can test.

And let me add that you decide when to stop the cycle, and usually after a few laps, the recovery goes horizontally, and then you burn it down.

Questions that are often asked

Will the critique cycle replace the design system?

No, they work out different jobs, the design system keeps up what you're going to produce; the review cycle monitors whether or not it fits the system, and the best result comes out when the two of them are together.

Can I use this when I write code?

Well, the logic is the same. It also works on the code side to put a freshly-connected review step independent of the breeding model. The difference is that there are already tests in the code, where the critic's job is largely tested.

How many critics is enough?

Three roles are a good balance in practice, two goes blind, five goes up, and the critics start repeating each other's identification.

How long should the rule file be?

It's not a target, it's a measure, you can re-establish the original by looking at the rules.

In a nutshell

The inconsistency is not because the model is incompetent, but because the evaluation is in the same place as the manufacturer. Separate the assessment from independent critics or by written rules, and your design becomes the system by accident. Just take the cost into account and save this heavy machine for the jobs you're going to use again.

Comments