gonzalo@flores — ~/libro/en/parte-2/05-higiene-sociotecnica sitio ↗
Acervo · Gonzalo Flores living digital book ES
Book contents You are here: Sociotechnical hygiene
II · The method
mature ~13 min read ch. 5/6 updated: 2026-09-23

Sociotechnical hygiene

Launch day has a deceptive way of looking like an ending. After weeks or months of work, the system responds, the accounts have been created and a first real transaction makes it through the new circuit without breaking. Screenshots are taken, someone congratulates the team and the calendar starts filling up with the next project. If there was training and a manual, the delivery is considered complete.

The life of a system, however, begins when the project that produced it ends. Until then it was used by people who knew its goals, could talk to the people who built it and paid extraordinary attention to every result. After launch it enters the ordinary workday. Someone in a hurry uses it, a piece of data arrives that was not in the examples, a supplier changes, a new person joins and the process that seemed stable bends to an emergency. That is when one finds out whether an artifact or a capability was built.

Distribuidora Norte had managed to put its alerts for at-risk customers into production. Salespeople checked the tool and the branches, after several arguments, shared a definition of “last purchase”. Six months later the system was still running and its availability indicators were flawless. Even so, something was breaking. On Fridays, when deliveries piled up, a parallel spreadsheet reappeared. Some alerts were closed without comment. Two new salespeople had learned the procedure by watching their colleagues, and each had learned it differently. From the infrastructure’s point of view, everything worked. From the work’s point of view, the bridge was starting to come apart.

What this book calls sociotechnical hygiene is the set of practices of care that prevents that decay. The word hygiene is deliberately ordinary. It does not describe an exceptional audit or a rescue operation, but modest, repeated habits: checking whether definitions still hold, listening to the workarounds the team invented, testing a backup, talking through a model error, updating who can authorize an exception. None of these actions has the glamour of a launch. Together they determine how long its value lasts.

The wear that does not show on the technical dashboard

Systems degrade in two ways. The first is familiar: a dependency ages, capacity runs short, a certificate expires, latency rises or the distribution of the data shifts. Technical teams have instruments to observe much of that decay. The second is harder to detect, because it happens in the relationship between the tool and the organization.

A category changes meaning in practice but not in the glossary. A permission granted as an exception becomes routine. The owner of an indicator changes jobs and no one inherits the obligation to review it. A screen requires recording a sequence that no longer matches the work, so people fill in the system afterwards, as a clerical chore. Each mismatch is small. Their accumulation produces sociotechnical debt: the future cost of having left the link between the human part and the technical part untended.

Debt is recognized by its interest payments. More sources have to be reconciled, more questions answered, and it has to be explained again and again why the dashboard contradicts experience. The team starts trusting names instead of processes: “ask Laura, she knows which number is the right one”. The people in charge feel that the system asks them for work and no longer gives help back. In the end comes the phrase that is usually read as resistance: “we used to do it faster”. Sometimes it is an idealization of the past. Other times, an accurate diagnosis.

Hygiene begins by taking that phrase seriously. It does not promise to go back, but it does not dismiss it either. It asks which task grew, which exception was lost and which part of the benefit stopped being visible to the person carrying the daily load.

Care is not an endless maintenance contract

A support contract may be necessary, but on its own it is not enough. Support repairs faults and answers questions. Hygiene builds, inside the organization, the capacity to notice that the system and the work are drifting apart.

That takes concrete owners. “The IT department” cannot own the definition of an active customer, the criterion for rejecting a recommendation or the way a commercial exception is handled. It can look after the infrastructure and make changes easier, but the meaning belongs to the business, and the consequences to those who do the work. Each routine needs an identifiable person, a reasonable frequency and a criterion that says when to act.

Not every organization needs a formal committee. In a small business, a monthly forty-minute conversation between operations, administration and the person who looks after the data can achieve more than a governance structure copied from a multinational. What matters is that the conversation happens, looks at evidence and ends with decisions. What workarounds appeared? Which figure caused arguments? What changed in the process? Is there an alert everyone ignores? Is someone using another tool because the official one does not meet their need?

The answers should leave a brief trace. Living documentation does not try to record everything. It keeps what a newcomer would need to understand the system’s reasoning: definitions, owners, known limits, design decisions and the conditions that would require revisiting them. A huge document no one opens is less useful than a precise page that changes along with the process.

What shadow AI is saying

In many organizations, the first sign that there is a need for artificial intelligence does not arrive as a formal proposal. It arrives when someone discovers that an employee is using a public assistant on their own to summarize documents, draft replies or analyze a spreadsheet. They may have pasted internal information. Perhaps they knew the policy forbade it and chose not to say so.

The immediate reaction tends to swing between two extremes. One celebrates the initiative: at last people are innovating without waiting for permission. The other sees an offense that has to be shut down. Both miss part of the phenomenon. Shadow AI, the use of AI tools outside the channels the organization has approved, can expose data and produce results that cannot be audited. It also reveals a demand the official channels were not meeting.

The 2025 global study by the University of Melbourne and KPMG found that almost half of the people surveyed admitted to uses that went against their organization’s policies. IBM, in its report on the cost of data breaches from the same year, linked a high presence of shadow AI to extra costs of hundreds of thousands of dollars. The figures give the size of the risk, but they do not explain why it happens. For that, one has to look at the work.

Whoever pastes a contract into a public model to get a summary probably did not start the day intending to break a policy. They had to read fifty pages before a meeting and found a tool at hand. Whoever drafts replies with a personal account may be covering a demand that grew while the team did not. The behavior may be reckless and need a clear limit. At the same time, it points with remarkable precision to where there is a costly task and an expectation of help.

That is why shadow AI enters the diagnosis as a symptom. First, the real uses are reconstructed, without organizing a witch hunt. Then the data and the impact of each case are classified: correcting a public text is not the same as summarizing a medical record or recommending a credit decision. After that, governed paths are offered. Some uses can move to a tool with a contract and adequate controls; others will require anonymization, human review or a ban. In some cases it will turn out that the best answer is not AI at all, but simplifying the process that pushed people toward the shortcut.

A ban without an alternative only makes the practice hide better. Openness without rules normalizes the risk. Hygiene works in the less comfortable space between the two: it acknowledges the need, makes the cost visible and builds an option people can use without lying about what they do.

From pilot to retirement

Governance accompanies the whole lifecycle. An idea that is still exploratory does not need the same controls as a system in production, but it does need to know what conditions it will have to meet to move forward. Checkpoints between exploration, pilot and operation make it possible to stop a solution before enthusiasm turns into dependency.

In the pilot, the questions are whether the use case is well defined, what data it touches and how the result will be measured. Before going into production, owners, monitoring, appeal mechanisms and an incident plan are added. During operation, drift is watched. And when the system is no longer justified, it is retired deliberately: the necessary records are kept, access is revoked and people are accompanied toward the flow that replaces it.

Retirement tends to be forgotten because it contradicts the story of permanent growth. But a system no one dares to switch off keeps consuming budget, attention and risk surface. Hygiene includes checking whether the tool still solves the problem that gave rise to it. Keeping something out of inertia is also a form of debt.

When there are models, monitoring cannot be reduced to availability. A complaint classifier can keep responding and still start treating certain groups differently because the data or the usage changed. In high-impact systems, algorithmic assessment has to be a periodic practice: if the context changes, the launch-day report no longer describes the current risk.

The human behavior around the model also needs watching. An acceptance rate that is too high can be as worrying as one that is too low, because people may have stopped reviewing. A growing volume of manual corrections may indicate drift, but also that a case appeared that the design never anticipated. The numbers open the investigation; they do not close it.

The boring part that keeps the business running

Some practices only show their importance when they are missing. Backups are the classic example. Setting up an automatic copy brings peace of mind, but the only real test is being able to restore it. A backup that has never been tested is a hope, not a protection.

The same goes for two questions that infrastructure abbreviates as RTO and RPO: how long a service can be down, and how much data can be lost without serious harm. They are usually presented as technical acronyms, though they express business decisions. How many hours can a point of sale go without operating? How many orders can the team rebuild if the database fails? The answer should not be invented by whoever manages the server. It has to be agreed with the people who know what an outage costs.

The cloud does not remove these responsibilities either: it splits them between provider and client. An organization can outsource its infrastructure without outsourcing the duty to manage access, classify data, test recoveries and prepare an exit. If the provider changes its terms or raises its prices, portability stops being an architecture discussion and becomes bargaining power.

Connectivity deserves the same attention. From an office with stable internet it is easy to treat the network as a detail. At a point of sale, a warehouse or a service center, an outage immediately turns into queues, lost sales and people improvising. Designing for contingencies means knowing what it costs to be disconnected and deciding accordingly, not overbuilding.

After an incident, the same logic that reflective practice applies to projects holds: reconstructing which barriers gave way and which signals were available produces better controls; looking for a quick culprit produces silence.

Teaching the judgment

The usual handover shows screens and procedures. That is enough as long as everything happens as it did in the demo. Real learning is tested in the face of an exception. Whoever receives the system needs to know why a rule was chosen, where it stops holding and whom to ask when the context changes.

That transfer of judgment can take several forms: working together on real cases, documenting the hard decisions, rotating responsibilities and letting the internal team lead a review with support. The goal is for the team to be able to reason without the system’s designer, not to repeat their words.

Training also has to reach those who join later. If all the knowledge lives in the people who took part in the project, each new hire resets the risk. An organization with high turnover needs short, frequent learning paths built into the work. An intensive course taught once a year can be flawless and still not solve its problem.

Marking the change

Previous processes do not disappear because the new system is available. For a while they coexist in people’s heads, and in them there are skills, pride and also ways of protecting oneself. If the launch ridicules everything that came before, the message to those who sustained the work is that their experience lost its value overnight.

That is what transition rituals are for. A retrospective can acknowledge what the old process did well and which cost is no longer worth paying. That closure does not idealize the past; it incorporates what was learned. The opening, for its part, needs an explicit space for the first complaints. A remark made in the first week is usually a design input. Ignored for months, it turns into resentment and a workaround that is hard to undo.

Rituals do not require solemnity. They can be a short meeting, a demo led by the users themselves or the simple gesture of giving the new flow a name. What matters is recognizing that change also touches identities: who knows, who helps, who authorizes, who is exposed. Those positions are negotiated in practice, not on the org chart.

Knowing when to step back

The quality of the bridge shows in what happens when the person who built it is no longer there. If every decision goes back to the consultant, if no one dares to update a definition or if an incident can only be solved by calling the outsider who knows the history, the system works at the cost of a dependency.

Stepping back well is a stage of design. It means leaving owners, routines and memory behind; checking that the organization can run a review without help; agreeing when it is worth asking for an outside view and when it is not. The advisor role can continue through spaced-out audits, but it should not replace internal capacity.

Sociotechnical hygiene pursues an outcome that is barely visible: that the system stops needing heroics. That data keeps its meaning, that exceptions have a place, that newcomers can learn and that a problem reaches the conversation before it becomes a crisis. None of that makes for a memorable photo. It produces something better: a transformation that still makes sense when no one remembers launch day anymore.

See also: Concepts of our own · Psychology of adoption · Value metrics · The bridge applied to the public sector