エピソード

  • Disaggregated Inferencing: يعني إيه وليه شراكة AMD و Cerebras مهمة؟
    2026/09/19

    Send us something, Share your comments directly :)

    هل شراكة AMD و Cerebras هتغير شكل داتا سنترز الذكاء الاصطناعي؟

    في الحلقة دي، بنشرح الإعلان المهم اللي حصل في مؤتمر AMD Advancing AI 2026، وليه معمارية الـ Disaggregated Inference بقت أهم بكتير من فكرة إنك تزود كروت (GPUs) وخلاص.

    عملية الـ Inference في الـ LLMs مش خطوة واحدة؛ دي بتتقسم لمرحلتين مختلفين تماماً عن بعض في متطلبات الهاردوير:

    مرحلة الـ Prefill: ودي عملية معقدة حسابياً (Compute-bound) ومحتاجة قدرة معالجة ضخمة عشان تعالج الـ Input وتطلع أول Token.

    مرحلة الـ Decode: ودي معتمدة تماماً على سرعة الذاكرة (Memory-bandwidth bound) وزمن الاستجابة عشان تبدأ تطلع الـ Tokens ورا بعض.

    في الحلقة دي هنتكلم عن:

    ليه الـ Prefill والـ Decode كل مرحلة فيهم محتاجة نوع هاردوير مختلف تماماً؟

    شرح مبسط لمفاهيم زي الـ TTFT و ITL وإيه هو الـ KV Cache.

    ليه تشغيل المرحلتين على نفس الـ GPU Pool بيعمل عنق زجاجة (Bottleneck) ويهدر موارد السيرفرات؟

    إزاي بنفصل الـ Workloads باستخدام راك AMD Helios مع شريحة Cerebras WSE-3؟

    تحديات الفابريك والشبكات: يعني إيه Latency Tax وقت نقل الـ KV Cache بين الكلاسترز المختلفة؟

    بتشبيهات وأمثلة عملية من الواقع، هنشوف إزاي المعمارية دي شغالة، وليه مستقبل البنية التحتية للذكاء الاصطناعي مبقاش متوقف على عدد الـ GPUs، بل على أوبتيمايزيشن المسار الكامل (End-to-End) للـ Tokens.

    والسؤال الأهم: هل الـ Disaggregated Architecture هتكون هي المعيار الجديد، ولا تكلفة الشبكات والـ Latency هتعطل الفكرة؟

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    11 分
  • الاربعة الكبار!: تحليل لمنتجات اريستا و جونيبر و سيسكو و هواوي في ال AI Datacenter
    2026/06/01

    Send us something, Share your comments directly :)

    In this episode of TelcoBytes, we explore the AI-Ready Datacenter from a Network and Infrastructure perspective. As Artificial Intelligence continues to reshape the technological landscape, the underlying network must evolve to handle unprecedented bandwidth and latency demands.

    We provide a comprehensive technical comparison of the AI networking portfolios from the top four industry leaders:

    • Arista Networks: We discuss their dominance with Hyperscalers, heavy reliance on Merchant Silicon (Broadcom Tomahawk series), and innovations like the Distributed Etherlink Switch to optimize Job Completion Time (JCT).
    • HPE Juniper Networks: A deep dive into their hybrid silicon strategy leveraging both Express 5 (Custom Silicon) and Tomahawk. We cover their use of High Bandwidth Memory (HBM) for mitigating incast traffic and Apstra's intent-based, multi-vendor fabric management capabilities.
    • Cisco: An overview of their Full Stack approach, strategic hardware partnership with Nvidia, the Silicon One architecture, and their push for power efficiency using Linear Pluggable Optics (LPO).
    • Huawei: An analysis of their vertical integration and full-stack AI ecosystem, highlighting the CloudEngine series, NSLB (Network-Scale Load Balancing) for lossless performance, and advanced liquid cooling designs.


    Whether you are building a massive GPU cluster or optimizing enterprise AI workloads, understanding these network architectures is critical. Tune in to discover how to choose the right infrastructure for your AI deployments.


    Follow Us:

    https://www.linkedin.com/in/telco-bytes

    https://www.linkedin.com/in/ledeeb

    https://www.linkedin.com/in/bassem-aly

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    27 分
  • كنز التيلكو في عصر الـ Inferencing
    2026/02/24

    Send us something, Share your comments directly :)

    Welcome to a new episode of TelcoBytes! Today, we are mapping out the next 12 months in the world of AI-Ready Datacenters. After spending trillions on training massive Large Language Models (LLMs), the industry is aggressively pivoting towards Inferencing in 2026.


    In this discussion, we decode "Inference Economics" and explain why generating output tokens behaves so differently than traditional CPU workloads. We also explore the intense "Silicon War" where Hyperscalers like Google, Microsoft, and AWS are developing custom chips to challenge Nvidia's dominance, while AMD makes strategic plays to secure its market share.


    Finally, we highlight the Golden Opportunity for Telco providers. With centralized datacenters facing physical latency limits for critical use-cases (like robotics and autonomous vehicles) and strict data sovereignty regulations, Telcos are perfectly positioned to win. By transforming Central Offices and RAN sites into distributed Micro-Datacenters for Edge AI, telecom operators can move from merely providing "dumb pipes" to delivering fully hosted, ultra-low latency AI applications.

    Follow Us:

    LinkedIn (TelcoBytes): https://www.linkedin.com/in/telco-bytes

    LinkedIn (Mohamed Eldeeb): https://www.linkedin.com/in/ledeeb

    LinkedIn (Bassem Aly): https://www.linkedin.com/in/bassem-aly

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    57 分
  • The Datacenter in the GenAI Era: What Changed? (Part2 - Power & Cooling)
    2026/01/12

    Send us something, Share your comments directly :)

    In this episode of Telco Bytes, we continue our exploration of the AI-Ready Datacenter, shifting focus from networking to the critical physical infrastructure: Power and Cooling.

    As AI workloads demand unprecedented computational power, the traditional data center design is being challenged. We analyze a real-world case study involving Meta (Facebook), discussing why they had to halt and redesign a major facility to accommodate the power-hungry nature of modern GPUs like Nvidia's H100.

    Key takeaways from this episode:

    • The Sunk Cost Fallacy in Tech: Why tearing down a partially built facility was the right strategic move for Meta to achieve faster Time-to-Market.
    • Power Dynamics: A walkthrough of the power journey from high-voltage transmission lines to the substation, and finally to the rack.
    • The Critical Role of UPS: Beyond battery backup, we explain how UPS systems utilize double conversion (AC-DC-AC) to clean the power sine wave and protect sensitive AI hardware.
    • Scale: Why the definition of a "large" data center has shifted from 50MW to hundreds of Megawatts in the AI era.

    Join us as we bridge the gap between high-level strategy and low-level infrastructure engineering.

    Follow Us:

    https://www.linkedin.com/in/telco-bytes

    https://www.linkedin.com/in/ledeeb

    https://www.linkedin.com/in/bassem-aly

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    41 分
  • أمازون، جوجل، اوبن ايه اي : الهروب الكبير من إنفيديا!
    2025/12/21

    Send us something, Share your comments directly :)

    This episode covers AI Datacenter News from a Network & Infrastructure perspective. We break down why power delivery and interconnect are becoming the dominant constraints, and how hyperscalers are increasingly pursuing Non‑NVIDIA approaches (e.g., Trainium and TPU) to reduce dependency and improve efficiency at scale.

    Topics include the “AI Factories” concept (datacenters producing tokens, not just storing data), the reality of gigawatt-scale training, distributed datacenter ambitions and distance/latency tradeoffs, and the growing Gulf momentum: UAE (Khazna / Stargate) and Saudi (HUMAIN + AirTrunk).

    Follow Us: https://www.linkedin.com/in/telco-bytes
    https://www.linkedin.com/in/ledeeb
    https://www.linkedin.com/in/bassem-aly

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    58 分
  • The Datacenter in the GenAI Era: What Changed? (Part1 - AI Workloads & Networking)
    2025/11/27

    Send us something, Share your comments directly :)

    The Datacenter in the GenAI Era: What Changed?

    In this episode of TelcoBytes Arabic, we tackle the fundamental question: Why do we need AI-Ready Data Centers, and what has fundamentally changed in the GenAI era?

    We explore this question through three distinct perspectives:

    ━━━━━━━━━━━━━━━━━━━━

    PERSPECTIVE 1: Traditional vs AI Workloads

    We compare E-commerce architectures (like Amazon) with AI Training Clusters to understand the fundamental shift:

    - Traditional Datacenters: Loosely coupled microservices that scale independently
    - AI Clusters: Tightly coupled systems where 100,000 to 1,000,000 GPUs must work as a single unit
    - Scale difference: From thousands of servers to millions of GPUs
    - Performance metrics: Transactions per Second vs PetaFLOPS

    ━━━━━━━━━━━━━━━━━━━━

    PERSPECTIVE 2: Network Challenges in the AI Era

    The Surprising Reality: Approximately 2/3 of Job Completion Time in AI Training is wasted on the Network!

    Key Challenges Discussed:

    TAIL LATENCY PROBLEM
    - How the slowest single frame can stall millions of GPUs
    - The Butterfly Effect: 1-2 millisecond delay can cause hours of training delay
    - Synchronization barriers where all GPUs wait for the slowest one

    GO-BACK-N PROTOCOL
    - Why AI uses RDMA over Converged Ethernet (RoCE)
    - Packet loss catastrophe: Much worse than latency
    - How Go-Back-N retransmits entire windows when one frame is lost

    ELEPHANT FLOWS
    - Few massive flows (Terabytes) vs many small flows
    - Low entropy in traffic headers
    - Traffic polarization: All traffic on one link while others remain idle

    INCAST PROBLEM
    - Many-to-one communication patterns
    - Congestion hotspots in the fabric
    - Buffer overflow even with deep buffers

    ━━━━━━━━━━━━━━━━━━━━

    PERSPECTIVE 3: Power & Cooling Implications

    How AI infrastructure requirements transform datacenter design:
    - Significantly higher power density
    - New cooling requirements
    - Time-to-market vs cost trade-offs

    ━━━━━━━━━━━━━━━━━━━━

    KEY TAKEAWAY

    The Network isn't just a connection between servers—it's the true Backbone and Nervous System of AI Data Centers. That's why NVIDIA calls it the "AI Backbone": without optimized networking, even the most powerful GPUs cannot operate efficiently.

    All these challenges have solutions, which we'll explore in detail in upcoming episodes!

    ━━━━━━━━━━━━━━━━━━━━

    TOPICS COVERED

    AI-Ready Datacenter | GenAI Infrastructure | Network Architecture | GPU Training | Traditional vs AI Workloads | Tail Latency | Job Completion Time | Go-Back-N Protocol | RDMA | RoCE | Elephant Flows | Traffic Polarization | Incast Problem | ECMP Hashing | All-to-All Communication | NCCL | Collective Operations | Deep Learning Infrastructure | Spine-Leaf Architecture | Data Center Networking

    ━━━━━━━━━━━━━━━━━━━━

    FOLLOW US

    TelcoBytes: https://www.linkedin.com/in/telco-bytes
    Mohamed Ledeeb: https://www.linkedin.com/in/ledeeb
    Bassem Aly: https://www.linkedin.com/in/bassem-aly

    #AIDataCenter #GenAI #NetworkArchitecture #DeepLearning #GPUTraining #DataCenterNetworking #InfrastructureEngineering #TelcoBytes #ArabicPodcast

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    41 分
  • من الـPixels للـ Parameters: اساسيات الـ GPU و الـ Model Training
    2025/10/28

    Send us something, Share your comments directly :)

    اهلا في أولى حلقات سلسه عالم مراكز البيانات الجاهزة للذكاء الاصطناعي – AI Ready Data Centers
    قبل ما نتكلم عن تصميم الشبكات، الـ fabrics، والـ interconnects، لازم نفهم الأساس:
    ليه الـ GPU أهم من الـ CPU في عالم ال AI؟
    وإزاي بيحصل الـ Training جوا الـ Deep Neural Networks؟
    وليه الشركات كلها بتجري تبني Data Centers مخصوصة للذكاء الاصطناعي AI؟
    في الحلقة دي هنشرح المفاهيم دي بطريقة بسيطة وسلسة،
    من قصة Colossus لإيلون ماسك وال500 ألف GPU واللي بتبرز السباق المحموم لبناء ال AI-DCs،
    إلى مفاهيم زي الـ FLOPS والـ Precision،
    وشرح عملي لشبكات ال Neural Networks خطوة بخطوة واللي بتعتبر اساس ثوره ال AI الحديثه.


    في الحلقة دي هتتعلم:
    ليه الـ GPUs قلب ثورة الذكاء الاصطناعي
    الفرق بين AI / ML / DL بطريقة بسيطة
    إزاي بيتم تدريب النماذج خطوة بخطوة كاساسيات لأي مهندس شبكات أو مهتم ببنية الـ AI

    🎧 اسمع الحلقة وادعمنا بالاشتراك 🔔
    💬 قول لنا في الكومنتس:
    عايزنا نتكلم أكتر في أي جزء المرة الجاية؟


    📢 تابعنا على لينكدإن لمزيد من الحلقات والنقاشات التقنية:
    https://www.linkedin.com/in/telco-bytes

    https://www.linkedin.com/in/ledeeb

    https://www.linkedin.com/in/bassem-aly

    #TelcoBytes_بالعربي | #AIReadyDataCenter

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    58 分
  • البداية!
    2025/10/20

    أولاً، لازم نقولكم إنكم وحشتونا جداً!

    إحنا رجعنا تاني، والمرة دي بموضوع هو "حديث الساعة" في مجالنا: الـ AI-Ready Datacenter

    كلنا شايفين طفرة الـ AI، لكن إيه اللي بيحصل في "المطبخ" عشان كل ده يشتغل؟ إيه اللي محتاج يتغير في الداتا سنتر بتاعتنا عشان نستحمل الـ Workloads الجديدة دي؟

    في الحلقة دي، إحنا مش جايين نتكلم في كلام نظري. إحنا هنمسك الموضوع بشكل عملي جداً وهنجاوب على مجموعة مهمة من الاسئلة. الأسئلة دول هما بالظبط اللي أي حد فينا محتاج إجاباتهم عشان يفهم يعني إيه داتا سنتر جاهزة للـ AI، وإزاي يبتدي يبنيها صح من الألف للياء.

    💬 قولنا رأيك:

    لو الحلقة عجبتك، متنساش تعمل Subscribe وتفعّل الجرس 🔔 وقولنا في التعليقات: ايه اكتر موضوع او سؤال عايز تعرف اجابته؟

    يلا بينا؟

    Hosts:

    Mohamed Eldeeb
    https://www.linkedin.com/in/ledeeb


    Bassem Aly
    https://www.linkedin.com/in/bassem-aly

    Follow us on

    Apple Podcast
    Google Podcast
    YouTube Channel
    Spotify


    続きを読む 一部表示
    12 分