VIP 👤
🏠 الرئيسية
المقاييس
📊 جميع المقاييس 🦖 ديناصور v1 🦖 ديناصور v2 ✅ تطبيقات قائمة المهام 🎨 صفحات حرة إبداعية 🎯 FSACB - العرض النهائي 🌍 مقياس الترجمة
النماذج
🏆 أفضل 10 نماذج 🆓 نماذج مجانية 📋 جميع النماذج ⚙️ كيلو كود
الموارد
💬 مكتبة الأوامر 📖 قاموس الذكاء الاصطناعي 🔗 روابط مفيدة 🔌 واجهات API والموجّهات
advanced

Deep Learning Model Compression

#machine-learning #optimization #deep-learning

Optimize a neural network via pruning and quantization.

Provide a technical guide on reducing the inference latency of a large Transformer model (e.g., BERT-base) for deployment on edge devices. Your guide should detail the process of: 1. Structured vs. Unstructured pruning—explain the trade-offs in hardware compatibility. 2. Post-training quantization (PTQ) vs. Quantization-Aware Training (QAT)—provide specific scenarios where one is preferred over the other. 3. Knowledge Distillation—describe how to structure the loss function between teacher and student models. Include code snippets using PyTorch or TensorFlow to demonstrate the implementation of a custom pruning schedule.