01
KV Cache Hit Ratio 修正模型:从直觉到统一公式
ai-systems / llm-inference
llm-inferencekv-cachesimulatorprefill
+1
02
量化:从 Scale 到 W4A8 的完整坐标
ai-systems / llm-inference
llm-inferencequantizationprefilldecode
+1
03
Prefill Trace:Worker 供给、DSA/MLA 与 Chunked Prefill
ai-systems / llm-inference
llm-inferenceprefilltracemla
+2
04
GDN 与 Chunked Prefill:为什么 prepare_chunk_indices 会出现在 trace 里
ai-systems / llm-inference
llm-inferencegdnqwen3nextchunked-prefill
+3
05
Chunked Prefill 深入分析:调度、Chunk Size 与 Attention 形状
ai-systems / llm-inference
llm-inferencechunked-prefillschedulingprefill
+3
06
Causal Attention:为什么 KV hit 后 Attention 按 1 - h² 缩放
ai-systems / llm-inference
llm-inferenceattentionkv-cachesimulator
+1
07
Compute-bound vs Memory-bound:推理的两大瓶颈
ai-systems / llm-inference
LLMInferencePerformanceGPU
+3