LET'S CONNECT ↗
01 — ESSAY

WHAT INDUSTRY CAN TEACH IT OPERATIONS

Reliability begins before production — and with the people closest to the work.
12 MIN READ · 10 AUG 2026 · VINCENT DEFARGE · OPERATIONS LAB
Industrial operations environment
THE PREMISE

In a factory, no one would claim that operations begin only when a machine stops.

Reliability is prepared long before that moment. The production system is designed, standards are established, abnormal conditions are made visible and teams are equipped to react before a small deviation becomes a larger failure.

Yet in IT, Operations are still often positioned at the end of the lifecycle: Architecture designs, Engineering builds, and then the service enters production. From that point onward, Operations are expected to monitor, support and restore it.

This model asks Operations to manage consequences they had too little opportunity to prevent.

Reliable digital services require operational knowledge much earlier. Before monitoring can be meaningful, someone must define normal behaviour. Before an alert can help, teams must understand which deviation matters. Before recovery can be prepared, failure modes, dependencies and decision paths must be explicit.

This does not mean Operations should own Architecture or Engineering. It means these disciplines must design the service together, with operational reality represented from the beginning.

Industry cannot provide a model for IT to copy literally. A physical production line and a distributed digital service behave differently. But industry can teach us something fundamental about the conditions required for reliable work.

GOOD OPERATIONS ARE NOT SIMPLY MANAGED.
THEY ARE DESIGNED.
01 — SEEING THE SYSTEM

RELIABILITY STARTS WITH VISIBILITY.

During visits to manufacturing environments, I was struck by the effort devoted to making the production process understandable while it is running.

Cameras provide visibility at different stages of the line. Weight, dimensions and other characteristics are measured continuously. Controls are distributed throughout the process rather than concentrated only at the end.

The objective is not to collect as much data as possible. It is to recognise a meaningful deviation early enough to protect what comes next.

Monitoring often tells us whether a known technical condition has occurred: a server is unavailable, a threshold has been exceeded, a transaction has failed or an application is responding too slowly. These signals are necessary, but they do not describe the whole service.

A service can remain technically available while already deteriorating for its users. A workaround may consume hours of operational effort. A dependency may become increasingly unstable. Repeated friction may affect a business process without ever generating a major incident.

Technical telemetry alone may not reveal these conditions. Some are first detected by the people operating the service or by the users experiencing it.

01

TECHNICAL

How the system behaves.

02

BUSINESS

What the user experiences.

03

OPERATIONAL

What teams encounter while keeping it working.

Observability is not about watching everything. It is about detecting meaningful deviation early enough to act.

A signal creates value only when the organisation understands what normal and abnormal mean, knows who should respond and has the capacity to act. In that sense, observability is not only a property of the technology. It is also a property of the operating model.

02 — FROM LEAN TO LEADERSHIP

A TOOL IS NOT A MANAGEMENT SYSTEM.

My understanding of the relationship between operations, people and management was strongly influenced by Jean-Christophe Guérin, a former senior manufacturing leader at Michelin.

I met Jean-Christophe through a mentoring programme. My current manager had known him much earlier in her career: he had been her first manager at Michelin and had worked on the start-up of new factories. Through that connection, I attended his conference on operational excellence and the inversion of the managerial pyramid.

I expected to learn mainly about industrial methods. What stayed with me was a conception of leadership.

The people closest to the work usually encounter its weak signals first. They see the recurring friction, the imperfect standard, the workaround and the small deviation before these appear in a management indicator.

The manager's role is not to replace that knowledge. It is to create an environment in which it can surface, help remove obstacles that teams cannot remove alone and progressively increase their ability to act.

This is the practical meaning of the inverted pyramid: management supports the people who create and operate value.

Jean-Christophe presented a problem-escalation mechanism known as the WIN system. Employees could make difficulties visible through WIN cards: red when help was required, green when information needed to be shared.

The card was only the signal. The management commitment behind it was the real system: a problem raised from the field could not simply be ignored.

Lean is easily reduced to its visible tools — standards, boards, indicators, cards and routines. Those mechanisms work only when people feel safe to reveal abnormalities and when the organisation responds constructively.

Without a response, visual management becomes visual inventory: a growing display of known problems that everyone has learned to tolerate.
03 — A CHAIN OF HELP

WHAT WOULD WIN MEAN FOR IT?

We have begun experimenting with the WIN approach in an IT context at Michelin. At this stage, its purpose is straightforward: make visible the operational problems that the people encountering them cannot resolve alone.

The experiment led me to a broader hypothesis of my own.

Could the principle of a chain of help operate not only between teams and managers, but also between Operations, Engineering and Architecture throughout the life of a digital service?

This is an extension of the original management idea, not a claim that the industrial mechanism can be transferred unchanged.

An Operations support team should be able to diagnose and resolve as many situations as reasonably possible. Autonomy, however, does not mean isolation or unlimited responsibility. Some problems require access, engineering capability or product knowledge that the team does not possess. Others reveal a structural design decision that no operational procedure can compensate for sustainably.

In those situations, the team should not have to choose between accepting the problem and transferring ownership of it.

01

OPERATIONS

Detect and frame the problem. Resolve locally when knowledge, authority and tools allow it.

02

ENGINEERING

Remove the technical constraint while Operations provides evidence and validates the result.

03

ARCHITECTURE

Address structural causes in boundaries, dependencies, resilience choices or quality attributes.

04

RETURN

Translate the solution into better tooling, automation, standards, documentation or decision rights.

The fourth step is what turns support into learning. Without it, the same dependency on Engineering or Architecture will remain the next time the problem occurs.

If a WIN simply means “I cannot fix this, so I transfer it”, it becomes another L1-L2-L3 escalation process. Ownership moves from queue to queue, response times increase and the boundaries between teams become stronger.

ESCALATE THE OBSTACLE.
NOT THE OWNERSHIP.

The team receiving the WIN joins the problem-solving process with a capability the originating team does not have. The team that raised it retains the context, participates in validation and receives the resulting knowledge.

Success should not be measured by the number of WINs created or closed. The more important question is whether the organisation becomes more capable over time.

  • Are recurring obstacles permanently removed?
  • Do teams recover effort previously lost to workarounds?
  • Can Operations resolve more situations safely without escalation?
  • Do operational signals influence engineering priorities and architectural decisions?
  • Does each response reduce the probability or impact of recurrence?

The objective is not to create a better escalation mechanism. It is to make each request for help improve the operating system.

04 — SIGNAL QUALITY

NOT EVERY IMPORTANT TOPIC IS A WIN.

Introducing a mechanism is easier than changing the behaviour around it. People do not always feel comfortable raising problems to management. Reporting a difficulty may be interpreted as complaining, admitting weakness or criticising another team's work.

Another distortion can appear once the mechanism gains management attention. WINs may gradually become a list of strategic topics rather than concrete obstacles encountered where the work happens.

NOT A WIN“We need to improve the CMDB.”

A broad ambition without a field problem or actionable request.

A USEFUL WIN“The support team spends three hours every week correcting this information manually.”

The cause is known, but removing it requires a product change the team cannot implement.

A meaningful WIN should identify the observed problem, its impact, what has already been attempted and the precise capability, authority or decision required from elsewhere.

This should not become a bureaucratic form. The receiving side also needs an explicit commitment. A WIN may not guarantee an immediate solution, but it should guarantee acknowledgement, clarification, a decision and feedback.

Trust develops through consistency: acknowledge the signal, understand it at the source, decide what happens next and close the feedback loop.

05 — BUSINESS IMPACT

MAKE SILENT DEGRADATION VISIBLE.

Not every important operational problem causes a total outage. Some create delays, manual work, reduced performance or recurring disruption over several days. A service may continue to meet a technical availability target while the business quietly absorbs the cost of its degradation.

We introduced a declarative impact score, with points added for each day a business impact continued. The score was not intended to evaluate teams or create a precise financial calculation. Its purpose was to make persistent degradation visible enough to support prioritisation.

This created a common signal for a difficult conversation: should the organisation continue adding change, or temporarily redirect capacity towards reliability and resilience?

That decision should never be automated by a score. Declarative measures are subjective. A number can create false precision, hide important context or be inflated to gain priority.

The score must remain the beginning of a conversation, not the end of one.

When degradation remains invisible, project delivery appears to be the only activity creating value. Once its cumulative impact can be seen, reliability work can be discussed as an investment in business performance rather than a technical interruption to the roadmap.

06 — THE MODEL

TOWARDS AN OPERATIONAL LEARNING SYSTEM.

IT organisations already generate enormous volumes of information: alerts, incidents, problem records, tickets, post-mortems, user feedback, technical debt and operational workarounds.

The challenge is rarely the complete absence of signals. It is the fragmentation between them.

Technical alerts live in monitoring tools. Business impact is discussed elsewhere. Operational irritants remain in team conversations. Structural weaknesses enter engineering backlogs. Architectural decisions are documented in another place again.

Each fragment may be visible, while the system connecting them remains invisible.

  1. 01

    SENSE

    Detect technical, business and operational deviations as early as possible.

  2. 02

    QUALIFY

    Understand the problem at its source, its impact and what has already been attempted.

  3. 03

    RESPOND

    Resolve locally or activate the appropriate chain of help without transferring ownership.

  4. 04

    REMOVE

    Address the technical, organisational or architectural obstacle rather than only its symptom.

  5. 05

    LEARN

    Return knowledge and capability to the field, then verify whether recurrence and effort decrease.

Technical observability supports the first movement. WINs can connect the second and third. Engineering and architectural action enable the fourth. Leadership must protect the fifth, because learning requires time, feedback and decisions that may compete with delivery.

This model remains a hypothesis to be tested and refined, not a finished framework. But it offers a way to connect practices that are too often managed separately.

This is not manufacturing copied into IT. It is industrial discipline translated into a digital operating environment.

07 — THE LEADERSHIP LESSON

KNOWLEDGE IS DISTRIBUTED.

Jean-Christophe Guérin's influence on me extends beyond operational methods. It has shaped the way I think about my own role as a leader.

The complexity of modern services makes it impossible for one manager to hold all the relevant knowledge. The people closest to Operations often encounter reality first. They see weak signals, recurring workarounds and the distance between a process as designed and a service as actually operated.

Leadership does not lose value when it admits that knowledge is distributed. Its value changes.

The leader must create the conditions in which important signals can surface without fear, ensure that calls for help receive a response and make the trade-offs that individual teams cannot make alone. The longer-term responsibility is not merely to solve today's issue, but to ensure that the intervention leaves the organisation more capable tomorrow.

WHAT DID THE PEOPLE CLOSEST TO THE SERVICE SEE FIRST?

WHAT PREVENTED THEM FROM ACTING?

WHICH PART OF THE SYSTEM ALLOWED THE PROBLEM TO PERSIST?

WHAT SHOULD RETURN TO THE FIELD?

Operational excellence is not achieved by designing better processes around people. It is achieved by designing better systems with them.

That principle leads directly to the next question: what happens when the applications are available but the business flow is not? I explore it in The Service Is Not the Application.

“
The goal is not simply to restore the service. The system — technical, organisational and human — should be more observable, more capable and more reliable tomorrow than it is today.
VINCENT DEFARGE · OPERATIONS LAB
CONTINUE THE CONVERSATION

WHEN A PROBLEM SURFACES, WHAT HAPPENS NEXT?

Does your organisation transfer the ticket, or remove the obstacle and return capability to the team? I am continuing to test this question across Operations, Engineering and Architecture — and I would value your perspective.