A Few Things I Didn't Want to Leave Behind

A Few Things I Didn't Want to Leave Behind
View on original source
Category: SciTech
Share
Archive
Like
I haven't sent Daily Intelligence consistently over the last few days - or weeks for that matter, and I didn't want to simply resume today without acknowledging that. Rather than send several older editions separately, I thought it would be more useful to bring together the ideas that stayed with me and explain why I think they belong in the same conversation. There were three pieces I had been working through. One concerned resilience and the ability of services to continue when their underlying infrastructure changes. Another looked at what happens when a trusted control remains operational but can no longer be trusted. The third examined the increasingly long-term commitments organizations are making around AI infrastructure. At first, they seemed like separate subjects. The more I considered them, the more I saw a common management problem. We are building organizations around technologies that are becoming more dynamic, more consequential, and harder to predict. Yet many of our operating assumptions still depend on things remaining relatively stable. That is the thread I wanted to share. One of the earlier editions began with a fairly ordinary cloud infrastructure problem. Yahoo described how its analytics workloads can use a ranked set of acceptable virtual machine configurations rather than depending on one exact machine type being available. If preferred capacity is constrained, the platform can select an alternative and keep the workload moving. The technical detail is interesting, but the organizational implication is more important. For a long time, resilience has often been associated with preserving a preferred state. We design the environment, establish redundancy, and try to restore the original configuration when something fails. Modern infrastructure increasingly works differently. Compute is replaced, workloads move, capacity changes, and applications may continue operating while the underlying environment looks quite different from the one that existed an hour earlier. That suggests a different question for service owners. Instead of asking only how quickly we can restore the original component, we should ask how many acceptable ways the service has to continue delivering its outcome. The strongest service is not necessarily the one with the most stable components. It may be the one least dependent on any single component staying exactly where it is. This is not an argument against standardization or redundancy. It is an argument for understanding the difference between a preferred operating state and the only operating state. A service may have alternative capacity, a secondary provider, a degraded but acceptable customer journey, or a temporary manual process. Those alternatives are only useful if the organization has already considered the conditions under which they can be used, who has authority to activate them, and what happens to the service commitment while they are in place. That is where Service Continuity becomes more than a recovery document. It becomes part of service design. The second edition was prompted by Rapid7's research into a compromised HAProxy implementation. The load balancer could continue handling legitimate traffic while malicious functionality operated inside it. What interested me was the distinction between availability and trustworthiness. We have become very good at asking whether systems are running. We monitor response times, infrastructure health, application errors, transaction volumes and a growing number of other signals. But a system can be available and still behave incorrectly. An identity platform might grant access it should not grant. An automation might report success when the downstream business action failed. A monitoring platform might miss an important class of events. A security control might continue operating while its own integrity has been compromised. The more authority a technology has, the more consequential that distinction becomes. The evidence needed to trust a control should not always come from the control itself. I think this is particularly relevant as organizations introduce more autonomous systems. We are going to give agents permission to read information, modify records, invoke tools and eventually perform increasingly consequential actions. It will not always be enough for the agent to tell us that the task was completed. We will need to know whether the resulting business outcome was correct, whether the action remained within authorized boundaries, and whether independent evidence supports what the system reported. That is not simply an AI governance issue. It is an operational accountability issue. And it connects directly to Service Management because the service owner ultimately needs confidence in the outcome, not merely confidence that the underlying technology says it is healthy. The third edition came from a very different direction: the scale of AI infrastructure investment. The announcement of a planned one-gigawatt AI data-center campus in Telangana, India, with investment of up to approximately $7.4 billion, was a useful reminder that AI is becoming an increasingly physical and capital-intensive industry. We tend to experience AI as software. We type into an interface, call an API, or deploy an agent. Behind that experience are enormous commitments involving land, power, cooling, networking, chips, buildings and long-term commercial arrangements. The interesting tension is that the technology at the top of that stack may change many times during the life of the infrastructure underneath it. Models evolve. Hardware generations change. Inference becomes more efficient. Enterprise use cases mature. Today's assumptions about demand may look very different in a few years. Yet some of the commitments being made now will last much longer. We are making some of the longest-lived technology commitments in decades around one of the fastest-changing technologies we have ever managed. For most enterprise leaders, the lesson is not that they should be building data centers. It is that AI strategy increasingly requires a more deliberate understanding of reversibility. Which decisions can we change easily? Which create long-term dependencies? Where are we locking ourselves into a provider, architecture, contract or operating model before we fully understand how the technology will evolve? The same question applies at a much smaller scale. An enterprise may not be investing billions in infrastructure, but it may be signing long-term platform agreements, embedding AI into critical workflows, or building applications around assumptions that will be difficult to unwind later. The issue is not avoiding commitment. It is knowing which assumptions we are allowing those commitments to preserve. The more I look at these three subjects, the more I think they are all pointing toward the same leadership challenge. We need organizations that can remain dependable without requiring the technology underneath them to remain unchanged. That means resilience cannot depend entirely on restoring the original environment. Trust cannot depend entirely on a system's own account of its behavior. And investment decisions cannot assume that today's technology choices will remain optimal throughout the life of the commitment. There is a common principle underneath all three: The organization needs to preserve its obligations without becoming unnecessarily dependent on the permanence of its implementation choices. That is a Service Management problem in the best sense of the term. The service is the durable commitment. The technology is how we fulfill it. The service owner should be able to explain what outcome must remain intact, what alternatives are acceptable, what evidence demonstrates that the service is behaving correctly, and what decisions the organization needs to retain the ability to change. This is also where Enterprise Service Management becomes a useful strategic horizon. The same questions apply whether the service is delivered through IT, HR, Finance, Customer Operations, Procurement or a combination of human and automated capabilities. The underlying technology may be different. The management obligation is remarkably similar. If I were discussing this with a CIO or CTO, I would choose one critical service and ask the leadership team to examine it from three angles. First, how many acceptable ways can the service continue operating when its preferred technology is unavailable or changing? Second, how do we independently know that the systems and controls supporting it are behaving correctly? And third, which decisions have we made around that service that would be difficult to reverse if our assumptions changed? I suspect that conversation would reveal more about the organization's actual resilience than another review of infrastructure availability alone. It would also help distinguish between investments that make the service more dependable and investments that simply make the current implementation more permanent. Those are not always the same thing. I don't think the answer to all of this is to make technology less dynamic. Nor do I think organizations should become hesitant about investing in AI, automation or modern infrastructure. The opportunity is too significant for that. But as technology becomes more capable and more deeply embedded in the enterprise, I think leadership needs to become more deliberate about what should remain stable and what should be allowed to change. The service commitment should be stable. Accountability should be clear. The evidence supporting trust should be reliable. The implementation underneath those things should have room to evolve. The technology can change. The organization's responsibility for the outcome cannot. That is the idea I wanted to bring forward from the editions I missed. And it is probably the question I would leave with you this week:where has your organization made the current implementation so permanent that it has become difficult to improve the service itself? Cheers, —Waseem

(0)Comments

 

A note on cookies

Newshunt uses essential cookies to keep you signed in and to remember your language and country, so the site works the way you expect. With your permission, we'd also like to use analytics cookies to understand how people use Newshunt and improve it over time.

Accepting only affects analytics. To learn more, view our Privacy Policy or Terms & Conditions.