#Resilience

Every story tagged Resilience, curated for CIOs and IT leaders — ranked by source credibility, engagement, and freshness.

5 stories · open in the command center

  • Cloud & InfrastructureTechCrunchTim De Chant2m

    One fallen power line exposed a growing AI data center problem. Here’s how to fix it.

    A power line failure in Northern Virginia revealed a critical infrastructure vulnerability: 3+ gigawatts of AI data centers simultaneously disconnected from the grid, causing a cascading power imbalance that nearly triggered widespread outages and demonstrated how concentrated data center loads threaten grid stability. As data centers are projected to grow from 6% to 24% of PJM's grid demand by 2040, IT leaders must adopt resilient power solutions and coordinate with grid operators to implement ride-through capabilities rather than automatic disconnection during fluctuations. This emerging regulatory and operational requirement represents a fundamental shift in how data center infrastructure must be engineered, with early movers like ON.Energy demonstrating that sophisticated UPS systems and load management can transform data centers from grid liabilities into stabilizing assets.

  • Security & PrivacyCIO Online5m

    5 steps to secure your infrastructure in the frontier model era

    As AI frontier models accelerate vulnerability discovery faster than traditional remediation cycles, CIOs must shift from reactive security to proactive infrastructure resilience—treating uptime and continuous vulnerability discovery as core security controls rather than operational afterthoughts. Organizations deploying AI at scale need enterprise-grade infrastructure engineered for security by design, with multilayered controls, redundant systems, and rapid recovery capabilities to withstand AI-driven attacks that can chain multiple vulnerabilities in minutes. The strategic imperative is clear: infrastructure that supports mission-critical workloads today will determine whether AI deployments remain secure, compliant, and operationally viable in an environment where one billion AI agents are expected by 2029.

  • Cloud & InfrastructureHacker News3m

    American Express: Cell-Based Architecture for Resilient Payment Systems

    American Express has implemented a cell-based architecture for its core payments platform that isolates failures within defined boundaries, enabling independent operation of each cell without cascading failures across the system. This design approach reduces failure blast radius, maintains low latency through data locality, and improves scalability while supporting mission-critical transaction processing at global scale. For IT organizations, this architecture pattern demonstrates how to balance resilience and availability requirements against operational complexity through careful isolation of services, data, and infrastructure.

  • Security & PrivacyCIO Online6m

    Cybersecurity maturity is now a proof point for resilience

    Cybersecurity maturity has evolved from a defensive necessity into a critical business resilience indicator that reveals whether organizations can withstand scrutiny, change, and risk management at scale. CIOs and technology leaders must translate cyber risk visibility across systems, users, and vendors into executive-level business language, as gaps typically surface during organizational change, acquisitions, audits, and insurance reviews—indicating when informal practices must transition to formal, repeatable, and governable controls. The cost of cybersecurity failures has become too significant for boards to absorb, making cyber posture a direct reflection of operational maturity and organizational preparedness.

  • AI & MLHacker News3m

    Decoupled DiLoCo: Resilient, Distributed AI Training at Scale

    Google DeepMind's Decoupled DiLoCo introduces a breakthrough distributed AI training architecture that enables resilient, asynchronous model training across geographically dispersed data centers with 20x faster convergence and orders of magnitude lower bandwidth requirements than traditional methods. This innovation allows IT organizations to leverage heterogeneous hardware (mixing different GPU/TPU generations), isolate failures to prevent cascading outages, and convert stranded compute resources into productive capacity—fundamentally changing the economics and operational complexity of large-scale AI infrastructure. For CIOs, this represents a shift from tightly-coupled, single-site AI training dependencies to flexible, globally-distributed models that improve both cost efficiency and business continuity while maintaining performance parity.

Browse all tags