Building AI-Ready Cloud Infrastructure: Strategies for Scale, Observability, Resilience, and Efficiency
Hosted by IEEE Cleveland member Sasi Kiran Malladi
Artificial intelligence is reshaping cloud infrastructure at a pace that is challenging long-standing assumptions about scale, performance, operations, and control. As organizations deploy larger models, accelerate inference, and expand AI-enabled services across production environments, infrastructure teams must support new demands in compute, storage, networking, capacity planning, and reliability. At the same time, these environments must remain observable, resilient, efficient, and secure.
This session examines how AI workloads are changing cloud infrastructure design and operations, and why observability is becoming a foundational requirement rather than an optional capability. It will explore the operational pressures introduced by AI, including dynamic workload behavior, higher infrastructure intensity, faster failure propagation, and the need for deeper telemetry across increasingly distributed systems. The presentation will also discuss practical strategies for building AI-ready cloud platforms that improve visibility, strengthen resilience, optimize resource efficiency, and support long-term scalability in complex enterprise environments.
Date and Time
Location
Hosts
Registration
-
Add Event to Calendar
Loading virtual attendance info...
- 17415 Northwood Ave
- #201
- Lakewood, Ohio
- United States 44107
- Building: Icon Co Work
- Room Number: Networking Room
- Click here for Map
Speakers
Sasi
Biography:
Sasi Kiran Malladi is a technology executive and a Principal at Amazon with nearly two decades of experience driving large-scale cloud and enterprise transformations. He has worked across several critical industries, including Financial services, Healthcare, State and Federal government, enabling some of the world’s largest and most influential organizations in their cloud journeys. Prior to AWS, he served as Director of Solutions Architecture at CGI. He advises organizations on technology and AI strategy, cloud transformation, and operational resilience, helping organizations align innovation with business outcomes. His expertise is in building scalable, secure, and resilient systems through AI-powered observability and advanced cloud operations, focused on generative and agentic AI. He also serves as an advisory board member and mentor, guiding teams and leaders on adopting modern architectures and best practices in complex, regulated environments. Sasi is an active contributor to the technology community as a speaker, author, reviewer, and an executive board member. He has delivered more than 50 speaking engagements, blogs, workshops and webinars, sharing practical insights on cloud, observability, and AI, and continues to shape how organizations design reliable, transparent, and accountable systems at scale.
Email:
Agenda
1. Why AI is changing cloud infrastructure now
- How AI workloads differ from traditional enterprise application patterns
- Why model training, fine-tuning, retrieval, and inference create new operational demands
- The shift from cloud-native infrastructure to AI-ready infrastructure
2. The core infrastructure pressures introduced by AI
- Higher demand on compute, storage, and network architectures
- Capacity planning challenges in bursty and high-intensity environments
- Latency, throughput, and utilization tradeoffs in AI-enabled systems
3. Why observability becomes essential in AI-ready environments
- The limits of conventional monitoring for modern AI workloads
- The growing need for end-to-end telemetry across applications, platforms, and infrastructure
- Detecting performance degradation, resource contention, anomalous behavior, and failure patterns earlier
4. Building for resilience, efficiency, and operational control
- Designing cloud environments that can absorb dynamic workload shifts
- Improving operational response through visibility, automation, and prioritization
- Balancing scale, reliability, and cost efficiency in production AI environments
5. Practical architecture and operations strategies
- Foundational design principles for AI-ready cloud platforms
- Strengthening observability across distributed services, data flows, and infrastructure layers
- Approaches for improving scalability and long-term sustainability without losing governance and control
6. What comes next
- Emerging infrastructure patterns for AI-enabled enterprises
- How infrastructure teams should prepare for future operating models
- Key takeaways for architects, platform teams, and technology leaders