Hello, World!
Welcome to Fenriar Tech! This blog is dedicated to sharing in-depth technical articles, engineering deep dives, and system architecture explorations.
What to Expect
Here, you will find practical guides and theoretical breakdowns covering:
- Large Language Models (LLM): Instruction fine-tuning pipelines, post-training quantization techniques, and inference serving optimizations.
- CUDA & High-Performance Computing: Kernel profiling, Roofline modeling, warp-level primitives, and memory hierarchy optimization.
- Inference Engines: Deep dives into TensorRT, custom operator plugins, and model acceleration pipelines.
- Deep Learning from Scratch: Operator backpropagation derivations, attention mechanisms, and custom implementations.
Stay tuned for upcoming articles!