LET'S CONNECT ↗
02 — ESSAY

THE SERVICE IS NOT THE APPLICATION

Why criticality, observability and recovery must follow the business flow.
10 MIN READ · 25 AUG 2026 · VINCENT DEFARGE · OPERATIONS LAB
Two industrial environments connected by a single red operational flow
THE INCIDENT

In What Industry Can Teach IT Operations, I argued that good operations are not simply managed. They are designed.

A real incident helped me understand what that means.

Two critical applications were available. Both had support teams on call. But the integration connecting them did not.

The data flow stopped — and so did the workshop.

This happened in a large industrial organisation. One application captured information from an upstream process. Another supported a downstream workshop. An API carried the data required to keep the physical activity moving.

One night, that API failed.

The two applications had been classified as critical and had support teams on call. The integration service between them had a lower classification. It did not have equivalent out-of-hours support or resilience arrangements.

The information stopped moving. The physical items did not. They continued to accumulate while the workshop could no longer process them. The organisation had to wait until morning for the missing link to be restored.

What stayed with me was not that an API had failed. APIs fail.

We had protected both applications, but not the business flow between them.

At first, this looked like a classification problem. Later, I recognised the same pattern in monitoring after maintenance and rollback during failed releases.

A SOUND OPERATIONAL INTENTION HAD FAILED TO BECOME A SUSTAINABLE OPERATIONAL CAPABILITY.
01 — CRITICALITY

CRITICALITY STARTS WITH THE BUSINESS FLOW.

Organisations often assess criticality application by application.

It is understandable. Applications have owners, support models, service levels and technical teams. They are visible objects around which responsibilities can be organised.

But the business experiences outcomes, not applications: an order placed, production information received, a shipment prepared.

Each outcome crosses applications, interfaces, data stores, infrastructure and people. The Business Promise is delivered by the entire chain.

The Business Promise describes the outcome the organisation must protect. The Service Promise translates that expectation into what the digital service must consistently deliver.

When critical applications depend on an integration that lacks sufficient support, resilience or an alternative operating mode, the whole flow remains exposed. Protecting each end of the chain is not enough if the bridge between them has been designed for a different level of importance.

Criticality starts with the business flow and its tolerance for disruption. It must then propagate to the operational capabilities and dependencies required to keep the Service Promise.

This does not mean assigning the highest classification to every underlying component. A dependency should inherit the requirements of the business flow only when it is necessary to deliver the corresponding level of service.

Some capabilities may be essential to the nominal service. Others may be sufficient to maintain a minimum acceptable outcome during disruption.

The better question is therefore not simply: Is this component critical?

For which business outcome and operational state is this component required?

Otherwise, we create critical islands connected by fragile bridges.

02 — OBSERVABILITY

OUR USERS HAD BECOME PART OF THE MONITORING SYSTEM.

I have seen the same gap appear after security patching and other maintenance activities.

An application would be restarted. Servers were up, processes were running, services had restarted and every monitoring dashboard was green.

Yet some users could not connect. Others could access the application but could not save an order because it could no longer complete the transaction path to the database.

The components were available. The Service Promise was not being fulfilled. And the first alert came from a user.

As an immediate safeguard, people on call began performing manual tests after maintenance to verify that essential actions still worked. It was a sensible response, but it exposed an uncomfortable truth.

Our users had become the final layer of our monitoring system.

We could observe component health automatically, yet still needed a human being to prove that the service was usable.

Technical monitoring remains essential. But operational visibility must go further:

  • Are the dependencies communicating?
  • Can the critical business event complete from end to end?
  • Is business impact accumulating even though the technology appears healthy?
  • Do decision-makers have enough reliable information to act in time?

This is what Business Observability means in this context: not monitoring more components, but producing trustworthy evidence that meaningful business events are completing — and making their impact visible when they are not.

Manual verification can provide immediate protection. A sustainable model should progressively turn critical checks into repeatable and, where appropriate, automated end-to-end evidence.

Otherwise, everything can be green while the business has already stopped.

03 — RECOVERY

WE HAD A ROLLBACK REQUIREMENT, NOT A REVERSIBLE DESIGN.

The same distinction exists in recovery.

For each release, teams were asked to provide a way to roll back so that the service could be restored quickly if something went wrong.

The intention was right. But rollback was often too complex to perform. New columns had been introduced, schemas had changed, or data had been written in a structure that the previous version could no longer understand.

A simple operational action had become a technical and data-recovery problem precisely when time mattered most.

We had a rollback requirement. We had not designed a reversible change.

The issue was not necessarily that a team had ignored the requirement. By final release approval, the design choices that made rollback difficult had already been made.

A genuine recovery capability must influence schema evolution, version compatibility, data protection, restoration tests and the choice between rollback, roll-forward or a degraded operating mode.

It must also be demonstrated. The existence of a plan does not prove that the complete business path can be recovered under representative conditions.

Recovery cannot simply be requested at the end. It must be designed from the beginning, tested in reality and revalidated when architecture, data, dependencies or operating assumptions change.

04 — THE OPERATIONAL INTENT GAP

PRESERVE THE MEANING, NOT ONLY THE DOCUMENT.

These incidents were not the result of one careless team.

Each team could protect the object it owned and meet its local requirements. The failure appeared between those objects, disciplines or stages of the service lifecycle.

The intention was present: protect a critical activity, verify recovery after maintenance and restore the service quickly if a release failed.

What was missing was the continuity needed to transform that intention into something visible, testable and executable. I think of this as an operational intent gap.

BUSINESS PROMISE → BUSINESS PROCESS → SERVICE PROMISE →
OPERATIONAL CAPABILITIES → DEPENDENCIES → EVIDENCE → LEARNING.

Business tolerance provides the boundary for that chain: how much interruption or degradation can be absorbed before the Business Promise is compromised.

01

TRACEABILITY

Can a requirement or control be traced back to the business expectation that justified it?

02

CONSISTENCY

Do criticality, service levels, recovery, observability and support models still express the same intent?

03

PROPAGATION

When the service changes, can every affected capability, dependency, control and piece of evidence be revalidated?

04

FEEDBACK

When reality differs from the design, does the learning reach Business, Architecture, Engineering and Service Management?

A good decision may remain in a document, classification or checklist while its operational purpose can no longer be seen or executed. Preserving the document is not enough. The relationships and meaning must survive as the service evolves.

THE PROBLEM IS NOT A LACK OF FRAMEWORKS.

None of these ideas is new in isolation.

Business Architecture already works with business outcomes, capabilities and value streams. Business Continuity starts from impact and tolerated disruption. Continuous Architecture treats architecture as an evolving set of decisions and explicitly considers operability, resilience and feedback. ITIL provides a holistic Service Value System, value streams, practices and continual improvement.

DevOps and DORA research address deployability, database change management, fast feedback and shared responsibility. SRE translates user expectations into measurable reliability objectives and promotes symptom-oriented, user-facing monitoring. Observability makes end-to-end behaviour across distributed systems visible.

The problem is not that these disciplines are incomplete or wrong. It is that organisations often implement them through different teams, artefacts, data models and governance mechanisms. Each discipline can work locally while the shared intent gradually fragments between them.

Organisations do not lack frameworks. They lose continuity between them.

The question is not which framework should own the whole chain. It is whether they continue to describe, support and improve the same promise as the service changes.

05 — OPERATIONS BY DESIGN

SIX QUESTIONS TO ASK WHILE CHOICES ARE STILL REVERSIBLE.

Closing this gap does not make Operations responsible for everything or require another approval gate.

It requires the right questions to be asked while important choices are still reversible:

  1. 01

    THE PROMISE

    What business outcome must remain true, and what Service Promise supports it?

  2. 02

    THE TOLERANCE

    How much disruption can the business tolerate, and what minimum level of service must continue?

  3. 03

    THE PATH

    Which operational capabilities and dependencies deliver that outcome in nominal and degraded states?

  4. 04

    THE EVIDENCE

    What proves that the critical business event completes and that the resulting business state can be trusted?

  5. 05

    THE RECOVERY

    Have recovery and data mechanisms been demonstrated, and what changes would require revalidation?

  6. 06

    THE DECISION

    Who has the information, authority and support needed to act — and how will operational learning update the design?

Applied to the workshop incident, these questions would have made the integration visible as part of the critical business flow. They would also have forced an explicit discussion about its support model, resilience, detection, alternatives and evidence before the night of the failure.

The business defines the outcome and its tolerance for disruption. Architecture makes structure, quality attributes and dependencies explicit. Engineering turns decisions into working technology. Service Management translates and governs service expectations across the lifecycle. Operations contributes knowledge about failure, detection, support, recovery and the reality of running the service.

This is the research space I am exploring through Operational Architecture: not a replacement for these disciplines, and not another gate, but a way to make their transformations and relationships explicit so that operational intent remains traceable, coherent and able to evolve.

This is Operations by Design.

06 — THE TAKEAWAY

THE APPLICATIONS WERE AVAILABLE. THE BUSINESS FLOW WAS NOT.

The incident mattered because the criticality of the Business Promise had been lost somewhere between the flow itself and our models of applications, support and resilience.

Good operations do not begin when a service enters production.

They begin when the Business Promise is translated into a Service Promise, operational capabilities, dependencies, controls and evidence — and when operational learning can travel back in the other direction.

A service is not operational because every component is green.

“
It is operational when the Service Promise is being fulfilled — and resilient when the organisation can demonstrate that the Business Promise remains protected through disruption.
VINCENT DEFARGE · OPERATIONS LAB
CONTINUE THE CONVERSATION

WHAT DOES YOUR ORGANISATION PROTECT: THE APPLICATION, OR THE BUSINESS FLOW?

I am continuing to explore how operational intent can remain visible across Business, Architecture, Engineering, Service Management and Operations — and I would value your perspective.