<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>BlackMesa Labs</title>
    <subtitle>Engineering notes: what we built, what we broke, what we took away.</subtitle>
    <link href="https://mesa.black/en/feed.xml" rel="self"/>
    <link href="https://mesa.black/en/"/>
    <updated>2026-10-04T00:00:00+00:00</updated>
    <id>https://mesa.black/en/</id>
        <entry>
        <title>Our domain layer contains nothing but exceptions. On purpose.</title>
        <link href="https://mesa.black/en/our-domain-layer-is-only-exceptions/"/>
        <id>https://mesa.black/en/our-domain-layer-is-only-exceptions/</id>
        <updated>2026-10-04T00:00:00+00:00</updated>
        <summary>Nine contexts, twenty-nine commands, a single query handler, zero ports. What we kept of DDD and hexagonal architecture, what we turned down, and the five places infrastructure crosses the boundary anyway. This is not purist DDD, it is time to market.</summary>        <content type="html">&lt;p&gt;Articles about hexagonal architecture almost always show the same diagram: a circle in the middle, ports around it, adapters outside, and an arrow pointing inward. None of them ever shows what the &lt;code&gt;Domain/&lt;/code&gt; folder actually holds six months later.&lt;/p&gt;
&lt;p&gt;Here is ours. Nine contexts, and twelve domain files in total: nine exceptions, two enums, one role. Zero interfaces. Zero aggregates. Zero value objects. Calling that a business layer would be a lie, and that is precisely what this piece is about.&lt;/p&gt;
&lt;h2&gt;What we kept: the C in CQRS, not the Q&lt;/h2&gt;
&lt;p&gt;The bus configuration declares three channels with explicit semantics:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-yaml&quot;&gt;default_bus: command.bus
buses:
    command.bus:            # a command mutates state, exactly one handler
        default_middleware: { enabled: true, allow_no_handlers: false }
    query.bus:              # a query returns a value, a single handler
        default_middleware: { enabled: true, allow_no_handlers: false }
    event.bus:              # a domain event: 0..n subscribers
        default_middleware: { enabled: true, allow_no_handlers: true }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The actual count today: &lt;strong&gt;twenty-nine command handlers, one query handler.&lt;/strong&gt; The query bus exists, it is configured, and it is very nearly empty.&lt;/p&gt;
&lt;p&gt;That is not migration debt, it is a conclusion. A command earns its ceremony because it brings three things we did not have: a name for an intent (&lt;code&gt;ChangeCompanySubscriptionPlan&lt;/code&gt; is not &lt;code&gt;setSubscriptionPlan&lt;/code&gt;), a guarantee that exactly one place executes it — &lt;code&gt;allow_no_handlers: false&lt;/code&gt; fails at boot, not in production — and an obvious transaction boundary, the handler’s.&lt;/p&gt;
&lt;p&gt;A read brings none of that. Its shape is dictated by the screen that displays it: this page needs these seven columns, joined this way, sorted like that. Putting that on a bus adds no rule, adds a layer, and moves the SQL from one file to another. So we read through Doctrine repositories, directly, without apologising for it.&lt;/p&gt;
&lt;p&gt;One detail matters more than it looks: &lt;code&gt;default_bus: command.bus&lt;/code&gt;. A bare &lt;code&gt;dispatch()&lt;/code&gt; is a command. The default is the channel that mutates state — the one thing you never want travelling down the wrong channel by accident.&lt;/p&gt;
&lt;h2&gt;What we turned down: dependency inversion&lt;/h2&gt;
&lt;p&gt;This is the heart of the hexagon in the literature: the domain declares interfaces, infrastructure implements them, the dependency arrow points inward. We did not do it. Here is a full handler, uncut:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-php&quot;&gt;#[AsMessageHandler(bus: &#039;command.bus&#039;)]
final readonly class ChangeCompanySubscriptionPlanHandler
{
    public function __construct(
        private EntityManagerInterface $entityManager,
        private LoggerInterface $logger,
    ) {}

    public function __invoke(ChangeCompanySubscriptionPlanCommand $command): void
    {
        $company = $this-&amp;gt;entityManager-&amp;gt;find(Company::class, Ulid::fromString($command-&amp;gt;companyId));
        if (!$company instanceof Company) {
            throw CompanyNotFoundException::withId($command-&amp;gt;companyId);
        }
        // …
    }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;EntityManagerInterface&lt;/code&gt; as a direct dependency. Out of thirty-eight handlers, &lt;strong&gt;twenty-four&lt;/strong&gt; depend on it. The &lt;code&gt;Company&lt;/code&gt; entity comes from &lt;code&gt;App\Entity&lt;/code&gt;, shared by every context. There is no abstract repository, no port, no domain model distinct from the table.&lt;/p&gt;
&lt;p&gt;Two remarks on that example, because they say more than the diagram does.&lt;/p&gt;
&lt;p&gt;First, the command carries &lt;code&gt;string $companyId&lt;/code&gt;, not a &lt;code&gt;CompanyId&lt;/code&gt; value object. That is not laziness: the message has to be serialisable for asynchronous transport. The bus boundary dictates the shape of the message — infrastructure setting the signature of what we call the domain, from the very first line.&lt;/p&gt;
&lt;p&gt;Second, &lt;code&gt;Ulid::fromString(...)&lt;/code&gt;. The domain identifier exists in two forms depending on whether you are carrying it or querying with it, and that conversion is a storage constraint, not a business one.&lt;/p&gt;
&lt;h2&gt;What the boundary buys anyway&lt;/h2&gt;
&lt;p&gt;One thing, exactly one, and it is worth its price: &lt;strong&gt;coupling between contexts became visible in the import list.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The handler above lives in &lt;code&gt;App\Billing&lt;/code&gt;. It imports &lt;code&gt;App\Company\Domain\Exception\CompanyNotFoundException&lt;/code&gt;. That single import says billing depends on the Company context, and it says so at the top of the file, in one line, with no diagram to keep up to date. The cross-context dependency graph is a &lt;code&gt;grep&lt;/code&gt; over &lt;code&gt;use&lt;/code&gt; statements.&lt;/p&gt;
&lt;p&gt;That is modest. Compared with a codebase where everything lives in &lt;code&gt;App\Service&lt;/code&gt;, it is the difference between “we assume it’s coupled” and “here is exactly where”. That benefit is what we bought, and nothing else.&lt;/p&gt;
&lt;h2&gt;The five places infrastructure crosses&lt;/h2&gt;
&lt;p&gt;Here is what no diagram shows: the places where Doctrine decides the shape of a business operation. These are real leaks, all of them in the code today.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. The business rule written in two dialects.&lt;/strong&gt; Our policy for keeping internal companies out of commercial figures is one predicate. It exists in two versions, because some read paths are native SQL for performance and others are DQL:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-php&quot;&gt;public static function sql(string $alias = &#039;c&#039;): string  // native queries
{ return ($alias !== &#039;&#039; ? $alias.&#039;.&#039; : &#039;&#039;).&#039;excluded_from_stats = false&#039;; }

public static function dql(string $alias = &#039;c&#039;): string  // ORM queries
{ return $alias.&#039;.excludedFromStats = false&#039;; }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;One rule, two spellings, because of the query language. The divergence is a single character — &lt;code&gt;excluded_from_stats&lt;/code&gt; against &lt;code&gt;excludedFromStats&lt;/code&gt; — so it is invisible to a quick read, and no test would fail if one of the two drifted.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. The ambient filter.&lt;/strong&gt; Doctrine’s soft-delete filter is mutable global state. A query looking for a unique identifier finds nothing while the unique index still holds it. The way out is to disable the filter and put it back:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-php&quot;&gt;$filters = $this-&amp;gt;em-&amp;gt;getFilters();
$wasEnabled = $filters-&amp;gt;isEnabled(&#039;softdeleteable&#039;);
if ($wasEnabled) { $filters-&amp;gt;disable(&#039;softdeleteable&#039;); }
$existing = $repo-&amp;gt;findOneBy([&#039;slug&#039; =&amp;gt; $slug]);
if ($wasEnabled) { $filters-&amp;gt;enable(&#039;softdeleteable&#039;); }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A “business query” whose result depends on ambient state is not a function. This is not a Doctrine defect — the filter does exactly what we asked of it — but it forbids reasoning about application code without knowing which mode it runs in.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Write ordering inside the unit of work.&lt;/strong&gt; Replacing a piece of content’s translations looks like one operation: drop the old ones, add the new ones, save. Within a single &lt;code&gt;flush()&lt;/code&gt;, Doctrine runs &lt;code&gt;INSERT&lt;/code&gt;s before orphan &lt;code&gt;DELETE&lt;/code&gt;s, and the unique constraint on &lt;code&gt;(feedback_id, locale)&lt;/code&gt; breaks. The operation has to be split into two successive saves.&lt;/p&gt;
&lt;p&gt;Put differently: the shape of the application-level operation is not set by the business but by the ORM’s internal scheduling. No port would have protected us, because the constraint lives neither in the domain nor in the adapter — it lives in the execution order that connects them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. The join alias that truncates in silence.&lt;/strong&gt; Filtering on an association already joined with &lt;code&gt;fetch&lt;/code&gt; also restricts the hydrated collection: you think you are filtering rows, you are amputating the returned object. Filtering needs its own second join:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-php&quot;&gt;$qb-&amp;gt;innerJoin(&#039;f.industries&#039;, &#039;i_filter&#039;)      // alias distinct from the fetch one
   -&amp;gt;andWhere(&#039;i_filter.id IN (:industries)&#039;);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The bug is invisible: the query returns the right entities, with incomplete collections. That is infrastructure semantics, right in the middle of what reads like a search rule.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;5. The identifier’s type.&lt;/strong&gt; Identifiers travel as ULIDs, compare as RFC 4122 in queries, and forgetting the conversion yields an empty result rather than an error. A port would have moved the conversion; it would not have removed it.&lt;/p&gt;
&lt;h2&gt;The migration rule&lt;/h2&gt;
&lt;p&gt;Which leaves the question that kills rewrites: what do we do with the existing code, written as classic Symfony services?&lt;/p&gt;
&lt;p&gt;Nothing, as long as nobody touches it. The policy is written down: &lt;strong&gt;CQRS for new writes and new reads; existing code migrates on demand only.&lt;/strong&gt; Never a wholesale rewrite in the name of consistency.&lt;/p&gt;
&lt;p&gt;Because consistency is not a business outcome. Rewriting a service that works so it resembles its neighbours produces risk without producing value — and it is exactly the kind of work that justifies itself indefinitely, because its stopping criterion is aesthetic.&lt;/p&gt;
&lt;h2&gt;This is not purist DDD, it is time to market&lt;/h2&gt;
&lt;p&gt;The real reason has to be named, because everything above reads differently once it is on the table.&lt;/p&gt;
&lt;p&gt;We have already written here that &lt;a href=&quot;/en/the-code-we-dont-write/&quot;&gt;the cheapest feature is the one you don’t build&lt;/a&gt;. A port is a feature. So is an aggregate, so is a value object, so is an anti-corruption layer. Each one costs writing, costs reading for whoever arrives next, and costs upkeep for as long as the code lives. An abstraction is not free because it is immaterial.&lt;/p&gt;
&lt;p&gt;So the useful question is not “is this correct DDD”. It is: what does this abstraction buy today, and what is it merely postponing? A repository behind an interface buys the ability to change storage — we are not leaving PostgreSQL. It also buys tests without a database, and that one is a real benefit: it is the single argument that may yet make us pay the bill.&lt;/p&gt;
&lt;p&gt;What we did instead fits in one sentence: ship, and keep only the boundaries that pay for themselves. The command bus costs three files and helps the same day. Ports cost a whole layer and may help, later, a team that does not exist yet.&lt;/p&gt;
&lt;p&gt;The risk in that position is well known, and writing it down is part of the price: it is indistinguishable from laziness at a glance. The difference comes down to one detail — we can name what we did not build, and why.&lt;/p&gt;
&lt;p&gt;And we own it. These are our choices, not accidents we would discover on a re-read. We correct them when we can and when we have the time: the day a friction becomes real and measurable, not the day an architecture article explains that they are incorrect. A debt that is written down, dated and argued is not a denied debt — it is the only kind you are able to repay at the right moment.&lt;/p&gt;
&lt;h2&gt;What it produced&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Nine named contexts&lt;/strong&gt;, whose mutual coupling is read in the imports rather than in an out-of-date diagram.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Twenty-nine commands&lt;/strong&gt;, each with an intent name, a single handler guaranteed at boot, and an obvious transaction boundary.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One query handler&lt;/strong&gt;, deliberately: reads go through repositories.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Twelve domain files&lt;/strong&gt; — nine exceptions, two enums, one role — which is to say no business rule genuinely isolated from Doctrine.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zero rewrites&lt;/strong&gt; of the existing classic code, which keeps working alongside.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Good practices&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Treat every abstraction as a feature: it has to say what it buys today, not what it would allow one day.&lt;/li&gt;
&lt;li&gt;Adopt the pieces separately. A command bus delivers something without the rest; ports, aggregates and value objects are distinct purchases, each with its own bill.&lt;/li&gt;
&lt;li&gt;Pick the strictest default: the default bus is the one that mutates state, and a missing handler fails the boot rather than production.&lt;/li&gt;
&lt;li&gt;Use namespaces as a coupling detector rather than a barrier: they protect nothing, they make things visible.&lt;/li&gt;
&lt;li&gt;Write the migration rule down, with its stopping criterion. Without one, “harmonising the architecture” is infinite work.&lt;/li&gt;
&lt;li&gt;Look at what &lt;code&gt;Domain/&lt;/code&gt; actually contains before claiming you do DDD. The count is instructive.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Watch-outs&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The folder is not the boundary.&lt;/strong&gt; Creating &lt;code&gt;Domain/&lt;/code&gt;, &lt;code&gt;Application/&lt;/code&gt;, &lt;code&gt;Infrastructure/&lt;/code&gt; gives a comforting sense of measurable progress — and measuring an architecture by its folder count is the surest way to get none of the guarantees they suggest.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Twenty-four of thirty-eight handlers depend on the &lt;code&gt;EntityManager&lt;/code&gt;&lt;/strong&gt;: none can be tested with an in-memory repository. The test suite needs a real database. That is a cost, it is accepted, it is not free.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shared Doctrine entities are the real coupling&lt;/strong&gt;, and it is invisible in the folder tree. The day two contexts want the same table to diverge, that is the wall we will hit — not the absence of ports.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A business rule duplicated across two query dialects will eventually drift&lt;/strong&gt;, and today nothing would catch it. That is the most concrete debt in everything above.&lt;/li&gt;
&lt;/ul&gt;
</content>
    </entry>
        <entry>
        <title>We were our own best customer. That was a bug.</title>
        <link href="https://mesa.black/en/our-own-best-customer/"/>
        <id>https://mesa.black/en/our-own-best-customer/</id>
        <updated>2026-09-30T00:00:00+00:00</updated>
        <summary>Our company publishes on our platform — and was counting itself in our revenue, our funnel and our follow-up lists. How we took it out of the statistics without taking it off the site.</summary>        <content type="html">&lt;p&gt;Everyone tells you to use your own product. Nobody warns you about the side effect: the moment your company has an account, it enters your figures. Not in a footnote — in the revenue, in the conversion rate, in the list of customers to chase.&lt;/p&gt;
&lt;p&gt;Our case: BlackMesa publishes real case studies on &lt;a href=&quot;https://showmetherex.com/en/&quot;&gt;Show me the REX&lt;/a&gt;. The account is real, the plan is real, the articles get read. What is not real is the revenue — we do not invoice ourselves. So a zero-euro customer sat inside our MRR, inflated the plan breakdown, and filled a slot in the conversion funnel without ever having converted anything.&lt;/p&gt;
&lt;p&gt;The obvious fix — “exclude our own company” — is wrong. It assumes the problem is the company. The problem is somewhere else.&lt;/p&gt;
&lt;h2&gt;It isn’t “who to exclude”, it’s “what does this number answer”&lt;/h2&gt;
&lt;p&gt;The reflex is to look for a list of entities to leave out. The right question is asked metric by metric: which question is this number the answer to?&lt;/p&gt;
&lt;p&gt;A view counter answers “how many people read this text”. The answer is the same whether the reader came from our offices or from anywhere else: the page was served, it was read, the content exists. Excluding our views would make that number wrong.&lt;/p&gt;
&lt;p&gt;Monthly revenue answers “how much do our customers pay us”. Our own company is not a customer. Its presence makes that number wrong.&lt;/p&gt;
&lt;p&gt;Both metrics look at the same database, sometimes at the same row, and need opposite rules. That is not an exception to handle, it is the rule: &lt;strong&gt;the unit of exclusion is not the data, it is the question being asked.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The case that settles it: a view that counts and does not count&lt;/h2&gt;
&lt;p&gt;The clearest example is also the most uncomfortable, because it rules out every implementation shortcut.&lt;/p&gt;
&lt;p&gt;On the platform, a case-study view feeds two distinct things. On one side, audience: the counter shown on the article, the dashboard total, the traffic curve. On the other, a commercial signal: “which companies have read your case studies”, used to identify contacts worth following up.&lt;/p&gt;
&lt;p&gt;Same event, same record. In the first group, a read from our own offices counts — somebody really did read it. In the second, it must not count at all: we are not a prospect to call back, and an internal reader sitting in a lead list is a sales action triggered for nothing.&lt;/p&gt;
&lt;p&gt;So there is no global filter to install once at the door. Every query has to know which family it belongs to.&lt;/p&gt;
&lt;h2&gt;A flag in the database, not a constant in the code&lt;/h2&gt;
&lt;p&gt;First version, the quickest one: a hardcoded list of company names. It lasted an hour. A constant forces a deployment to reclassify a company, and more importantly it lies about the nature of the information: “this company is not a customer” is business data — it changes, it gets decided, and it has to be visible to whoever administers the accounts.&lt;/p&gt;
&lt;p&gt;So it became a checkbox on the company form, with its help text spelled out, because an option whose scope nobody understands ends up ticked at random:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Our own companies publish and stay visible on the public site, but count in no commercial or marketing figure: MRR, plans, funnel, upsell/churn, leads, digests. Views and visits, however, still count.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The list of names did not disappear: it bootstraps a freshly installed environment so it never starts with polluted metrics. But it no longer decides anything.&lt;/p&gt;
&lt;p&gt;The predicate itself lives in one small class, in two flavours — one for ORM queries, one for raw SQL. That sounds trivial; it is what makes the rule auditable. Finding every place that applies it became a text search, and checking it against the list of commercial figures takes a minute by eye.&lt;/p&gt;
&lt;h2&gt;Verify rather than re-read&lt;/h2&gt;
&lt;p&gt;Re-reading your own exclusion code proves nothing: you read back what you believe you wrote. The only verification worth the name is to toggle the checkbox and compare the dashboards before and after, figure by figure.&lt;/p&gt;
&lt;p&gt;About thirty values moved. Seven stayed strictly identical — and that was the expected result: they are the audience counters and the published-content volumes. A number that moves when it shouldn’t is a bug; a number that stays put when it should have moved is an omission. Without that pass, we would have caught neither.&lt;/p&gt;
&lt;p&gt;The final shape: one flag, one policy class, thirty-four predicates across eight files — MRR, plans, funnel, upsell, churn risk, follow-ups, lead intelligence, monthly digests, coupon expiry, nurturing.&lt;/p&gt;
&lt;p&gt;That last one deserves a mention. Without the exclusion, our own company sat in the automated re-engagement sequence. We would have been sending ourselves our own win-back e-mails.&lt;/p&gt;
&lt;h2&gt;What it produced&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;An MRR with no zero-euro customer in it&lt;/strong&gt;, and a conversion rate that no longer counts our own sign-up in its denominator.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;About thirty corrected values&lt;/strong&gt; across the commercial dashboards, seven deliberately unchanged.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zero change on the public site&lt;/strong&gt;: the case studies stay published, visible, filterable, and their views still count.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One rule written in one place&lt;/strong&gt;, with its reason, rather than an &lt;code&gt;AND&lt;/code&gt; copied from query to query.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A checkbox&lt;/strong&gt; that reclassifies a company without a deployment.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Good practices&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Classify each metric before coding the filter: audience measurement, or commercial signal. The answer sets the rule, and it is not the same for two numbers drawn from the same table.&lt;/li&gt;
&lt;li&gt;Put membership in the database, not in a constant: it is business data, it changes and it gets decided.&lt;/li&gt;
&lt;li&gt;Keep the predicate in a single class even if it fits on one line: the point is not reuse, it is being able to find every call site.&lt;/li&gt;
&lt;li&gt;Verify by toggling: comparing figures before and after is the only way to tell an omission from a decision.&lt;/li&gt;
&lt;li&gt;Write the scope into the admin interface, right where the box gets ticked.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Watch-outs&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Nothing stops the next query from forgetting the predicate.&lt;/strong&gt; The rule is held by code review, not by a test, and that is the main weakness of the setup today.&lt;/li&gt;
&lt;li&gt;An exclusion that is too broad is as wrong as no exclusion at all: pulling our views out of the audience counter would have produced a lie in the other direction.&lt;/li&gt;
&lt;li&gt;The day an internal company becomes a real paying customer, the box has to be unticked — and nobody will remind you. That decision is human, and it needs a scheduled moment where it gets revisited.&lt;/li&gt;
&lt;li&gt;Saying publicly that your statistics exclude your own data costs nothing, and beats having someone else discover it for you.&lt;/li&gt;
&lt;/ul&gt;
</content>
    </entry>
        <entry>
        <title>The cheapest feature is the one you don&#039;t build</title>
        <link href="https://mesa.black/en/the-code-we-dont-write/"/>
        <id>https://mesa.black/en/the-code-we-dont-write/</id>
        <updated>2026-09-27T00:00:00+00:00</updated>
        <summary>Three signals that say &#039;do not write this code&#039; — and what they saved us in a single day.</summary>        <content type="html">&lt;p&gt;Digital sobriety is almost always framed as an optimisation problem: lighter images, better
caching, a cleaner region. All true, and all marginal. The real waste sits elsewhere, and
nobody measures it: &lt;strong&gt;code written for nothing&lt;/strong&gt;. A migration to redo, a tool bought then
dropped, a feature shipped that nobody opens. Which leaves the question we had never put
into words: how do you &lt;em&gt;decide&lt;/em&gt; not to build something?&lt;/p&gt;
&lt;p&gt;The report gives the &lt;strong&gt;three signals&lt;/strong&gt; we settled on, and what each one saved over one
ordinary working day: a broken feature nobody complained about — information about the
feature, not about the bug; 1,400 translations we did not write because no reader could
reach them; and a monitoring gap that turned out to be a rehearsal gap, not a tooling gap.&lt;/p&gt;
&lt;p&gt;It also says what those signals &lt;strong&gt;do not&lt;/strong&gt; mean. “Nobody needs it” is the perfect excuse for
doing nothing, and a rule that can only say no has stopped being a rule. Two guardrails keep
it honest — including this one: an unused feature is sometimes badly exposed rather than
useless.&lt;/p&gt;
</content>
    </entry>
        <entry>
        <title>Upgrading the OS under Docker: what is still coupled, and what nothing validates</title>
        <link href="https://mesa.black/en/upgrading-the-os-under-docker/"/>
        <id>https://mesa.black/en/upgrading-the-os-under-docker/</id>
        <updated>2026-09-27T00:00:00+00:00</updated>
        <summary>The application lives in an image, so the distribution cannot touch it. Four coupling points remain — and one of them sits in a blind spot no pipeline covers.</summary>        <content type="html">&lt;p&gt;&lt;a href=&quot;https://showmetherex.com/en/&quot;&gt;Show me the REX&lt;/a&gt; runs on a single host: one Postgres, one Redis, one front proxy, and the application deployed blue-green — two identical instances, “blue” and “green”, of which only one serves traffic at a time. A new version is started on the idle one, and once it answers correctly the proxy is switched over. At the time of writing, that host serves 52 published case studies and a little over 16,000 cumulative views.&lt;/p&gt;
&lt;p&gt;This arrangement protects a release: if the new version misbehaves, switching back takes a second and nobody notices. It protects nothing at all against the machine itself, since both instances sit on that same host — stop it and you stop them both.&lt;/p&gt;
&lt;p&gt;The single host is a deliberate choice, not an oversight: no redundancy until traffic justifies it, and the threshold for revisiting it is written down — a thousand visits a day. Below that, a second machine costs far more in complexity (replication, failover, consistency) than it returns in availability, and complexity nobody needs yet is the most expensive thing you can build.&lt;/p&gt;
&lt;p&gt;Owning that choice creates an obligation, though. If you accept that a reboot takes the service down, you owe yourself an exact figure for what it costs. That is the part we got wrong that night, and we come back to it at the end.&lt;/p&gt;
&lt;p&gt;We moved that host from one Ubuntu release to the next. The operation went through — and that is not the interesting part. The interesting part is that containerisation has narrowed what such an upgrade can still break down to a very short list, and that we found exactly the item on it that nothing in our chain was watching.&lt;/p&gt;
&lt;h2&gt;What containerisation actually decoupled&lt;/h2&gt;
&lt;p&gt;The application is immune to a distribution upgrade, and it is worth being precise about why: it does not use anything from the host. Its PHP, its extensions, its system libraries, its CA bundle all ship inside the image. The host’s &lt;code&gt;ca-certificates&lt;/code&gt; can be replaced wholesale — our outbound calls to VIES, the payment provider and object storage are unaffected, because they never read that store.&lt;/p&gt;
&lt;p&gt;This is why teams have come to treat an OS upgrade on a container host as routine. It is very nearly true, and it is the “very nearly” that costs.&lt;/p&gt;
&lt;h2&gt;The four coupling points that remain&lt;/h2&gt;
&lt;p&gt;Once you have moved everything you can into images, what still belongs to the host is a short list — and it is the same list on every containerised server:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;The kernel.&lt;/strong&gt; Containers share it. A major jump changes the ground under a database far more than under a web process.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Docker daemon.&lt;/strong&gt; It is not in your image; it is a host package, installed from a third-party repository.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The apt sources for that daemon.&lt;/strong&gt; The thing that decides whether the previous point ever gets a security patch again.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The restart contract.&lt;/strong&gt; Which containers come back on their own after a boot, and which deliberately do not — the idle instance, for one, must stay down, or two versions of the application would fight over the same database.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Nothing else really matters. That list is short enough to be checked by hand, before and after — which is precisely what makes not checking it inexcusable.&lt;/p&gt;
&lt;h2&gt;The only lasting damage landed on point three&lt;/h2&gt;
&lt;p&gt;A release upgrade disables or removes third-party apt sources. This is not a bug: those sources are built for the release you are leaving, and keeping them enabled across the jump is how you break a system. The tool is right to do it, and it says so.&lt;/p&gt;
&lt;p&gt;What it does not do is tell you afterwards. Our Docker source was gone. Nothing broke — the daemon kept running, the containers with it, the site served. But the installed package had been built for the previous distribution, and there was no longer any repository able to replace it. &lt;strong&gt;The failure mode is not a service that stops; it is a package manager that becomes authoritative and empty at the same time.&lt;/strong&gt; It reports that everything is up to date, and it is telling the truth about a universe it can no longer see.&lt;/p&gt;
&lt;p&gt;Restoring the source is four lines and no restart. It immediately revealed seven minor versions of drift accumulated in silence. Nothing would ever have raised a hand: not the daemon, which works; not the monitoring we do not have; not &lt;code&gt;apt&lt;/code&gt;, which had nothing left to compare against.&lt;/p&gt;
&lt;h2&gt;The blind spot: the compose files&lt;/h2&gt;
&lt;p&gt;Then the structural finding, which is the one worth taking away.&lt;/p&gt;
&lt;p&gt;Our database service had no restart policy. After a reboot, every other container would have come back and Postgres would not. The bug is trivial — one missing line. Its lifespan is not: it had been there for months, and it could only manifest on a host reboot, which had not happened in that whole time.&lt;/p&gt;
&lt;p&gt;The real question is not how it got written. It is why nothing caught it. And the answer generalises well beyond our setup: &lt;strong&gt;the compose files are the only production configuration that nothing owns.&lt;/strong&gt; They are not baked into the image, so the build never sees them. No test exercises them, because tests run against the application, not against the host’s topology. And our deploy job does not even copy them — it opens an SSH session and runs a script. They are edited in the repository, applied by hand, and validated by nothing.&lt;/p&gt;
&lt;p&gt;In a chain that is otherwise fully automated — tests, static analysis, image build, zero-downtime deploy — that is where a defect can sleep indefinitely. Not in the code the pipeline reads twenty times a day, but in the handful of YAML lines it never opens.&lt;/p&gt;
&lt;h2&gt;What the kernel could have broken, and what we did not verify&lt;/h2&gt;
&lt;p&gt;A major kernel jump under a containerised Postgres deserves more than “it came back up”. Three questions are worth asking, and we can only answer the first two:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;The data directory.&lt;/em&gt; It lives in a named volume, on the same filesystem, with the same storage driver. Nothing moved — and that is the reason the upgrade was survivable at all, not luck.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;The daemon’s confinement profiles.&lt;/em&gt; Seccomp and AppArmor defaults ship with the Docker package, not with the distribution, which is why the containers found the same environment on the other side.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Durability semantics.&lt;/em&gt; Whether a new kernel changes anything for Postgres under our mount options is a question we did not answer. It held, which proves nothing. We are writing it down rather than claiming a verification we never ran.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The measurement already existed&lt;/h2&gt;
&lt;p&gt;Which leaves the obligation set out at the start: knowing the real cost of the outage we chose to accept. We stated a figure by reading container logs — the gap between a process’s last error and the next one’s start. That number describes the host, not the visitors. It leaves out the shutdown before it and the warm-up after, and it can be off by a factor of three.&lt;/p&gt;
&lt;p&gt;Meanwhile all our traffic goes through a CDN, which had already recorded, request by request, exactly what people got during that window. &lt;strong&gt;The measurement we thought we lacked had been taken for us, by an appliance we had been paying for and never queried.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The reflex to fix is not “add instrumentation”. It is to inventory what already measures before adding anything: the CDN, the reverse proxy’s own logs, the provider’s console. Our procedural slips that night — a guessed container name, shell variables that do not survive an SSH reconnection, a grep loose enough to report 41 lines for 2 errors — all belong to the same family, and they are worth exactly one sentence: a runbook must resolve what it needs, never freeze it. Which instance is live is the textbook case: it changes at every deploy.&lt;/p&gt;
&lt;h2&gt;Why we publish all of this — and the one thing we hold back&lt;/h2&gt;
&lt;p&gt;A report is only useful if it is specific, so this one carries the commands, the failure modes and the gaps. Which raises a fair question: is publishing all that not handing an attacker a map?&lt;/p&gt;
&lt;p&gt;Our rule fits in one sentence. &lt;strong&gt;The problem is never naming a version. It is naming a version you are still vulnerable on.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;“We were on version X, it cost us this, it is patched” is ordinary post-mortem practice. The same text published before the fix is a weakness that is still true with the target attached — a case study is signed, it names a company and a domain. So: fix first, tell afterwards, and generalise whatever teaches nothing. The &lt;em&gt;gap&lt;/em&gt; — seven versions of silent drift — is the lesson; our host’s exact version string is not. What never gets published is the other category, the part that cannot be learned from but can be copied: host names, paths, the chain of keys that decrypts a backup. The line is not “sensitive versus harmless” — it is &lt;strong&gt;does this teach, or does this open?&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What it gave us&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Four coupling points&lt;/strong&gt; isolated between a containerised application and its host: kernel, daemon, the daemon’s apt sources, restart contract. Short enough to check by hand, before and after.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A frozen update path restored&lt;/strong&gt;: the daemon had drifted seven minor versions behind, with nothing able to report it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A structural blind spot named&lt;/strong&gt;: the compose files, the only production configuration that no build, no test and no deploy job ever reads.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A measurement recovered rather than built&lt;/strong&gt;: the CDN had already recorded what visitors saw, for free.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Good practices&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Write down your coupling points once. On a container host there are four of them, they are always the same, and checking them takes ten minutes.&lt;/li&gt;
&lt;li&gt;After a release upgrade, re-check the third-party sources first: a package manager with nothing left to compare against reports that everything is up to date, and it is not lying.&lt;/li&gt;
&lt;li&gt;Look for what your automation does not own. In a fully automated chain, the defect that survives is in the file the pipeline never opens.&lt;/li&gt;
&lt;li&gt;Inventory what already measures — CDN, proxy, provider console — before adding instrumentation. The figure you are missing has often already been recorded.&lt;/li&gt;
&lt;li&gt;A runbook resolves, it never freezes: which instance is live, its container name, the service id — all of it is asked for at call time.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Watch-outs&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The dangerous failure is not the service that stops, it is the one that keeps working while losing its ability to be updated. It emits nothing.&lt;/li&gt;
&lt;li&gt;“It came back up” proves nothing about durability. Write down the questions you did not answer rather than claiming a verification you never ran.&lt;/li&gt;
&lt;li&gt;A defect that can only manifest on a host reboot has the lifespan of the interval between two reboots. On a well-behaved server, that is months.&lt;/li&gt;
&lt;li&gt;An internal log says when a process died, never when visitors stopped being served. Only a measurement taken in front of the stack says that.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fix before you tell: a report describing a weakness that is still open, signed with your own name, is not transparency — it is an instruction manual.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
</content>
    </entry>
        <entry>
        <title>Building a complete, production-ready SaaS with a coding agent</title>
        <link href="https://mesa.black/en/a-complete-saas-with-a-coding-agent/"/>
        <id>https://mesa.black/en/a-complete-saas-with-a-coding-agent/</id>
        <updated>2026-09-26T00:00:00+00:00</updated>
        <summary>Zero-downtime deploys, real payments, passwordless sign-in — paired with an AI.</summary>        <content type="html">&lt;p&gt;How we designed, secured and shipped an entire platform — blue-green switching with no
downtime, real direct debits, passkey sign-in, a trilingual site, encrypted backups — while
working alongside a coding agent.&lt;/p&gt;
&lt;p&gt;The report goes through the architecture we settled on, what the agent genuinely
accelerated, and above all the places where we had to take the wheel back.&lt;/p&gt;
</content>
    </entry>
        <entry>
        <title>Guaranteeing legitimate case studies without a moderation team</title>
        <link href="https://mesa.black/en/legitimacy-without-moderation/"/>
        <id>https://mesa.black/en/legitimacy-without-moderation/</id>
        <updated>2026-09-25T00:00:00+00:00</updated>
        <summary>Betting on verified identity rather than censorship — and the two bugs that made us doubt it.</summary>        <content type="html">&lt;p&gt;A case-study platform is worth exactly as much as the trust placed in what gets published
on it. One fake report and the credibility of the whole corpus wobbles. The question we
faced: how do you guarantee the legitimacy of every publication, at scale, without
standing up a moderation team — and without handing the triage to an AI, which judges the
form of a text rather than the legitimacy of whoever publishes it?&lt;/p&gt;
&lt;p&gt;Our answer replaces content moderation with &lt;strong&gt;sender verification&lt;/strong&gt;: no company, no
publication; a minimum of substance required; an intra-community VAT number checked
against the EU registry; full traceability.&lt;/p&gt;
&lt;p&gt;The report also tells the story of the &lt;strong&gt;two traps&lt;/strong&gt; we hit along the way, both from the
same family: a third-party service answering &lt;code&gt;200 OK&lt;/code&gt; without having been able to check
anything, then an “unknown” state that actually covered two opposite situations — “I could
not check right now” and “I will never be able to check this value”. The second one let any
string of characters through for months.&lt;/p&gt;
</content>
    </entry>
    </feed>
