About Om Chimurkar
Om Chimurkar is an AI researcher and machine learning systems engineer focused on the intersection of deep learning theory, hardware execution efficiency, and long-context transformer inference. His research addresses the acute memory wall in modern large language models, specifically how Key-Value (KV) cache expansion bottlenecks real-world token generation throughput and swells GPU VRAM allocations.
Through novel mathematical formulations—including anchor-stabilized low-rank singular value decomposition (SVD) deltas in DKV (Differential KV-Cache), rigorous method-agnostic benchmarking in CRBench, and mechanistic interpretability of test-time compute in SpiralState—Om creates rigorous, reproducible systems that enable frontier language models to reason over massive sequences with verified mathematical stability and reduced computational costs.
Peer-Reviewed Publications & Technical Reports
PEER-REVIEWED RESEARCH • SUBMITTED TO IEEE TETCI
1. DKV: Anchor + Low-Rank Differential KV-Cache Compression
Author: Om Chimurkar (Independent Research) • Venue: IEEE Transactions on Emerging Topics in Computational Intelligence (IEEE TETCI) — Under Review • DOI: 10.5281/zenodo.21539110 • Code: github.com/Omc12/Differential-KV
Abstract: A training-free KV-cache compression runtime combining exact anchors, rank-32 joint K|V truncated SVD deltas, and multi-signal exact residual tokens to preserve verbatim recall while achieving 1.44×–2.25× memory reduction and 1.72× faster prefill scaling at 64k context on Apple M3 (8.6GB) over memory-optimized dense baselines. Validated with zero recall collapse (ΔPPL = 0.00%) across 4k–64k horizons where standard FP16 crashes with Out-Of-Memory at 16k.
Key Metrics: 1.44×–2.25× Memory Savings • 1.72× Prefill Scaling • ΔPPL = 0.00% • Tested on Qwen2.5-1.5B (64k context) on Apple M3 Silicon
SYSTEMS BENCHMARK • SUBMITTED TO ASPLOS 2027
2. CRBench: A Method-Agnostic, Resource-Aware Evaluation Framework for Long-Context LLMs
Author: Om Chimurkar (Independent Research) • Venue: ASPLOS 2027 — Under Review (September Review Cycle) • DOI: 10.5281/zenodo.22085120 • Code: github.com/Omc12/CRBench
Abstract: Designed a query-level, method-agnostic evaluation framework measuring contextual capability retention relative to an identical dense FP16/BF16 reference while explicitly accounting for memory and system resources with unified scoring S_res = αQ + (1 − α)R_mem. Isolates true compression degradation from intrinsic base-model hallucinations and tracks physical OS Resident Set Size (RSS) against analytical allocations.
Key Metrics: 98.6% Normalized Quality Score • 63.8% OS Memory Reduction • 1.31× End-to-End Speedup • Validated on LLaMA-3.1-8B and Qwen2.5-1.5B
REASONING DYNAMICS • SUBMITTING TO NAACL 2027
3. SpiralState: Emotional State and Reasoning Effort in Language Models
Authors: Om Chimurkar, Dhruv Ramani • Venue: In Preparation for NAACL 2027 • Code: github.com/Omc12/SpiralState
Abstract: Controlled empirical study investigating how prompt affective framing and emotional preambles influence reasoning effort, Chain-of-Thought deliberation length, and serving-time compute costs in language models. Negative emotional framing induces up to +34.6% reasoning token expansion without accuracy gains, causing repetitive verification loops. Probing residual stream activations across 32 transformer layers achieves 91.4% separation accuracy along an affective vector at Layer 18.
Key Metrics: +34.6% Reasoning Token Inflation • 91.4% Layer 18 Probing Accuracy • Models Tested: Qwen3.5-9B, IBM Granite 4.2-8B, QwQ
DETERMINISTIC EVALUATION • RESEARCH ARCHIVE
4. Deterministic AGI Benchmarking Under Constructed Cognitive Load
Author: Om Chimurkar (Independent Research) • DOI: 10.5281/zenodo.19652443
Abstract: Evaluated 7 frontier models across 116 original constructed-language ('Glorbish') rule-learning tasks spanning Legal, Science, Social Norms, and Math/Logic domains with deterministic algorithmic verification, eliminating LLM judge agreement bias and uncovering an authentic 23.2 percentage-point performance spread across frontier models.
RETRIEVAL SYSTEMS • EMPIRICAL ABLATION
5. An Ablation Study of Retrieval Strategies for Stock-News RAG Systems
Author: Om Chimurkar (Independent Research) • DOI: 10.5281/zenodo.19086005
Abstract: Controlled empirical ablations comparing dense semantic, BM25 keyword, and Maximal Marginal Relevance (MMR) retrieval, showing hybrid retrieval improves recall coverage (+18.4%) while MMR increases contextual diversity (+26.1%) orthogonally to factual grounding.
Technical Skills & Systems Expertise
LLM & Deep Learning Systems: PyTorch, Hugging Face Transformers, LLM Inference Optimization, KV-Cache Compression (DKV), Attention Kernels, Model Quantization, Low-Rank Truncated SVD, RAG Architectures, Mechanistic Interpretability, Representation Engineering.
GPU Computing & Performance Profiling: CUDA, Triton, Apple MLX, NVIDIA Nsight Systems, Nsight Compute, GPU VRAM Profiling, Memory Optimization, Kernel Benchmarking, Linux Systems Programming.
Programming Languages: Python, C++, CUDA C++, JavaScript, TypeScript, SQL, Bash.
Frameworks & Tools: FastAPI, React, Node.js, Git, Docker, Vector Databases (FAISS, ChromaDB), Linux CLI, Firebase.