GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)
Independently published
Pages: 75, Paperback, Independently published
Compare prices (1 shop)
| shop | Price | Action |
|---|---|---|
|
|
7,48 GBP | Go to shop |
