Efficient Deployment of Embedded Foundation Models at the Edge
IEEE Computer Society Columbia Section
Technical Webinar Series
Efficient Deployment of Embedded Foundation Models at the Edge
Date and Time
Location
Hosts
Registration
-
Add Event to Calendar
Loading virtual attendance info...
Speakers
Mahsa
Efficient Deployment of Embedded Foundation Models at the Edge
Large language models (LLMs) and vision-language models (VLMs) are enabling a new generation of intelligent applications. However, deploying these models on resource-constrained edge devices remains a significant challenge due to their computational and memory requirements.
This webinar covers three of our recent research projects that explore how to make modern AI systems more efficient and practical for edge computing. It begins with LLMPi, which investigates techniques for optimizing large language models to achieve high-throughput inference on Raspberry Pi. Building on this foundation, PiCASo, our ACM GLSVLSI 2026 Best Paper Award-winning work, combines optimized LLMs, Whisper speech recognition, and text-to-speech into a fully offline, real-time, end-to-end conversational AI system running on Raspberry Pi 5. Finally, VLMath extends these ideas to multimodal AI by leveraging efficient vision-language models to provide pedagogically aligned math tutoring through efficient multimodal reasoning.
The talk also discusses practical techniques for deploying AI models on resource-constrained devices, including model quantization, hardware-aware optimization, and efficient system design, while highlighting the challenges and performance trade-offs encountered across these projects.
Biography:
Mahsa Ardakani is a Ph.D. student in the Department of Computer Science and Engineering at the University of South Carolina and a Graduate Research Assistant in the Intelligent Circuits, Architecture, and Systems (iCAS) Lab under the supervision of Dr. Ramtin Zand. Her research focuses on the efficient deployment of large language models (LLMs) and vision-language models (VLMs) for edge computing and hardware-aware AI systems. Her work aims to make modern AI practical on resource-constrained devices for applications including conversational AI, multimodal intelligence, and educational technologies. Her recent work received the Best Paper Award at ACM GLSVLSI 2026.
Address:United States
Agenda
| Time | Agenda |
|---|---|
| 0:00 – 0:05 | Welcome and Opening Remarks (IEEE Computer Society Columbia Section) |
| 0:05 – 0:10 | Speaker Introduction |
| 0:10 – 0:45 | Technical Presentation: Efficient Deployment of Embedded Foundation Models at the Edge |
| 0:45 – 0:55 | Live Q&A Session |
| 0:55 – 1:00 | Closing Remarks, Upcoming IEEE Events, and Networking |
Organized by:
IEEE Computer Society Columbia Section (CH03263)
IEEE Columbia Section | IEEE Region 3
Join us to explore the latest advances in efficient AI deployment, edge computing, and embedded foundation models.