adobe machine learning engineer interview: ml fundamentals and optimization
Interview Experience
Transformer / Attention Fundamentals Difference between encoder vs decoder transformers Bidirectional attention vs causal masked attention Semantic purpose: understanding vs autoregressive generation Encoder-decoder cross attention behavior Why encoder does not need KV cacheMulti-Head Attention Internals Full operation order inside MHA Q/K/V projections and tensor reshaping Attention score computation and softmax Why there are 4 linear projections total FLOPs analysis of self-attention...
Full Details
Transformer / Attention Fundamentals Difference between encoder vs decoder transformers Bidirectional attention vs causal masked attention Semantic purpose: understanding vs autoregressive generation Encoder-decoder cross attention behavior Why encoder does not need KV cacheMulti-Head Attention Internals Full operation order inside MHA Q/K/V projections and tensor reshaping Attention score computation and softmax Why there are 4 linear projections total FLOPs analysis of self-attention Why attention becomes O(S^2)KV Cache / Inference Optimization How KV cache works Why cache K/V but not Q Why KV cache matters for decoder-only models Why encoder architectures do not benefit much Memory growth tradeoffs of KV cacheTransformer Systems / Performance Why transformers are memory-bound Memory bandwidth vs compute bottlenecks Attention matrix materialization cost HFU / MFU concepts GPU utilization and profiling Kernel launch overhead Kernel fusion optimization CUDA graphs and fused kernelsFlashAttention / Attention Optimization What FlashAttention does internally Tiling and SRAM usage Online softmax Avoiding HBM materialization Reducing memory traffic Other attention optimizations: * MQA * GQA * sparse attention * linear attention * paged attention * quantized KV cacheDistributed Training / Infra HFU tracking during training Communication bottlenecks GPU profiling approaches Throughput optimization techniquesInference Optimization / Deployment TensorRT and inference compilers Operator fusion Graph optimization Mixed precision inference vLLM / paged attention serving End-to-end serving latency considerationsMixed Precision Training FP16 vs BF16 vs FP32 Why mixed precision speeds up training Tensor core acceleration Loss scaling FP32 master weights Why BF16 is preferred for LLMs Exponent range vs mantissa tradeoffLLM Training / Alignment / Evaluation DPO training behavior Teacher-student distillation pipelines Offline teacher generation + online student serving LLM-as-a-judge evaluation Structured evaluation rubrics Offline vs online metrics Reward hacking / evaluator bias Alignment drift and hallucination regressionProduction ML / Latency Real-time latency constraints Offline feature generation vs online serving Feature store lookup design Long-tail fallback systems Lightweight online student models* Cache hit rates and serving metrics
About This Question
This is a candidate experience report from a adobe interview for a mle role (newgrad level) during the phone screen round reported in 2026.