The tools hand you a logo in ninety seconds. That is the least valuable part of what you need.
A mark, a typographic pair, a colour set with defined roles, a photographic treatment, a tone of voice, and rules for how these behave when they meet each other. The mark is the smallest component. Generate a hundred logos and you still have no identity, because nothing tells the hundredth asset how to agree with the first.
This is why today’s test is the grid test and not do you like the logo. Tools today: Pomelli, Looka, OpenArt, Google Labs.
Image models do not remember yesterday’s output. Every generation starts from nothing, so identical intent produces divergent results unless you carry the constraints forward yourself. The fix is a written constraint block you paste into every generation: exact colour values, the lens and framing language, the light quality, the treatment, the mood adjectives, and what is banned. Write it once, reuse it for the whole set. This is the same discipline as a shot list, you are not being creative each time, you are executing a decision made once.
Models cannot hit an exact hex reliably, they approximate. For anything where brand colour is contractual, generate the imagery and set the exact colour in a real editor, and name the colour in words as well as value, deep petrol green steers where a hex code does not. Text inside generated images is still unreliable, especially in Arabic, where letterforms connect and most models trained on almost no correct examples. Generate the image, set the type properly in a design tool, and never ship generated Arabic lettering to a Gulf client without a native speaker reading it. Faces and hands remain the first place a viewer’s eye finds the seam.
Subject, what is in frame, described specifically, a woman is nothing, a woman in her fifties in a tailored abaya, mid conversation, hands visible, is a brief. Composition, shot size, angle, lens, because these models learned from captioned photography, lens language moves them more than people expect. Light, the single largest lever on mood and the one amateurs leave out, direction, quality, colour. Treatment, film stock, grain, contrast curve, grade, this is what makes six images feel like one campaign. Negatives are a fine tuning instrument, not a strategy.
Default outputs skew Western because the training data does. Ask for an office and you get Californian open plan, ask for a family home and you get suburban America. If you work here, specificity is your job: name the region, the architecture, the light quality of a Gulf afternoon, the dress, the materials. Then look hard for the details a model will get plausibly wrong, the cut of a thobe, the way a ghutra sits, the geometry of a mashrabiya. A client will spot these instantly and it will cost you the room.
Pomelli reading a client’s website is genuinely useful, and it is reading the symptoms of a brand, not the brand. It cannot know the founder’s reason for starting, the promise made to the first customer, the thing the company refuses to do. Those come from a conversation with a human being. The gap between what the tool extracted and what the brand guidelines say is a precise measurement of what the website fails to communicate, and that gap list is often worth more to the client than the assets.
NOVA reacts, nothing is scored, nothing is stored against you.
Invent a small business. Generate a logo set in Looka, then use OpenArt to produce six social posts that visibly belong to the same brand. Put all six in a grid and show someone. If they cannot tell it is one brand, tighten your constraints and go again.
Run Pomelli against a real client website and let it build the Business DNA profile. Compare what it extracted against the actual brand guidelines and note every place it got the brand wrong. That gap list is a genuinely useful deliverable. Then produce a campaign set and a brand book a designer would accept.
Day 6 in progress
Tomorrow, Day 7: sound and voice. Identical text in two voices is two different messages.