OM CHIMURKAR

AI Researcher & Machine Learning Systems Engineer

Undergraduate AI Researcher at Newton School of Technology, Rishihood University (B.Tech in Artificial Intelligence, GPA: 8.0 / 10.0). Specializing in LLM inference optimization, transformer memory systems (DKV), KV-cache compression, low-rank SVD factorizations, custom CUDA / Triton / Apple MLX kernels, and reasoning compute dynamics (SpiralState).

STATUS: TRANSMITTING ACTIVE SIGNALS
LOCATION: DELHI NCR, INDIA
GitHub (@Omc12) • LinkedIn (/in/om-chimurkar) • Google Scholar • ResearchGate • ORCID (0009-0004-0518-4598) • omchimurkar45@gmail.com

About Om Chimurkar

Om Chimurkar is an AI researcher and machine learning systems engineer focused on the intersection of deep learning theory, hardware execution efficiency, and long-context transformer inference. His research addresses the acute memory wall in modern large language models, specifically how Key-Value (KV) cache expansion bottlenecks real-world token generation throughput and swells GPU VRAM allocations.

Through novel mathematical formulations—including anchor-stabilized low-rank singular value decomposition (SVD) deltas in DKV (Differential KV-Cache), rigorous method-agnostic benchmarking in CRBench, and mechanistic interpretability of test-time compute in SpiralState—Om creates rigorous, reproducible systems that enable frontier language models to reason over massive sequences with verified mathematical stability and reduced computational costs.

Peer-Reviewed Publications & Technical Reports

PEER-REVIEWED RESEARCH • SUBMITTED TO IEEE TETCI

1. DKV: Anchor + Low-Rank Differential KV-Cache Compression

Author: Om Chimurkar (Independent Research) • Venue: IEEE Transactions on Emerging Topics in Computational Intelligence (IEEE TETCI) — Under Review • DOI: 10.5281/zenodo.21539110 • Code: github.com/Omc12/Differential-KV

Abstract: A training-free KV-cache compression runtime combining exact anchors, rank-32 joint K|V truncated SVD deltas, and multi-signal exact residual tokens to preserve verbatim recall while achieving 1.44×–2.25× memory reduction and 1.72× faster prefill scaling at 64k context on Apple M3 (8.6GB) over memory-optimized dense baselines. Validated with zero recall collapse (ΔPPL = 0.00%) across 4k–64k horizons where standard FP16 crashes with Out-Of-Memory at 16k.

Key Metrics: 1.44×–2.25× Memory Savings • 1.72× Prefill Scaling • ΔPPL = 0.00% • Tested on Qwen2.5-1.5B (64k context) on Apple M3 Silicon
SYSTEMS BENCHMARK • SUBMITTED TO ASPLOS 2027

2. CRBench: A Method-Agnostic, Resource-Aware Evaluation Framework for Long-Context LLMs

Author: Om Chimurkar (Independent Research) • Venue: ASPLOS 2027 — Under Review (September Review Cycle) • DOI: 10.5281/zenodo.22085120 • Code: github.com/Omc12/CRBench

Abstract: Designed a query-level, method-agnostic evaluation framework measuring contextual capability retention relative to an identical dense FP16/BF16 reference while explicitly accounting for memory and system resources with unified scoring S_res = αQ + (1 − α)R_mem. Isolates true compression degradation from intrinsic base-model hallucinations and tracks physical OS Resident Set Size (RSS) against analytical allocations.

Key Metrics: 98.6% Normalized Quality Score • 63.8% OS Memory Reduction • 1.31× End-to-End Speedup • Validated on LLaMA-3.1-8B and Qwen2.5-1.5B
REASONING DYNAMICS • SUBMITTING TO NAACL 2027

3. SpiralState: Emotional State and Reasoning Effort in Language Models

Authors: Om Chimurkar, Dhruv Ramani • Venue: In Preparation for NAACL 2027 • Code: github.com/Omc12/SpiralState

Abstract: Controlled empirical study investigating how prompt affective framing and emotional preambles influence reasoning effort, Chain-of-Thought deliberation length, and serving-time compute costs in language models. Negative emotional framing induces up to +34.6% reasoning token expansion without accuracy gains, causing repetitive verification loops. Probing residual stream activations across 32 transformer layers achieves 91.4% separation accuracy along an affective vector at Layer 18.

Key Metrics: +34.6% Reasoning Token Inflation • 91.4% Layer 18 Probing Accuracy • Models Tested: Qwen3.5-9B, IBM Granite 4.2-8B, QwQ
DETERMINISTIC EVALUATION • RESEARCH ARCHIVE

4. Deterministic AGI Benchmarking Under Constructed Cognitive Load

Author: Om Chimurkar (Independent Research) • DOI: 10.5281/zenodo.19652443

Abstract: Evaluated 7 frontier models across 116 original constructed-language ('Glorbish') rule-learning tasks spanning Legal, Science, Social Norms, and Math/Logic domains with deterministic algorithmic verification, eliminating LLM judge agreement bias and uncovering an authentic 23.2 percentage-point performance spread across frontier models.

RETRIEVAL SYSTEMS • EMPIRICAL ABLATION

5. An Ablation Study of Retrieval Strategies for Stock-News RAG Systems

Author: Om Chimurkar (Independent Research) • DOI: 10.5281/zenodo.19086005

Abstract: Controlled empirical ablations comparing dense semantic, BM25 keyword, and Maximal Marginal Relevance (MMR) retrieval, showing hybrid retrieval improves recall coverage (+18.4%) while MMR increases contextual diversity (+26.1%) orthogonally to factual grounding.

Key Systems Benchmarks & Empirical Results

2.25×
DKV Peak Memory Savings

Compresses KV cache footprint up to 2.25× on 64k sequences with zero loss of retrieval accuracy.

0.00%
ΔPPL Perplexity Drift

Exact anchor tokens prevent RoPE phase coordinate drift across extensive context horizons.

98.6%
CRBench Quality Score

Normalized capability retention while delivering 63.8% physical OS Resident Set Size (RSS) reduction.

+34.6%
Reasoning Token Inflation

Documented in SpiralState under negative affective prompt framing, inflating inference costs without accuracy gains.

Now: Current Focus, Roadmap & Lab Dispatches

[01] Active Engineering & Research

  • Preparing SpiralState Tests & Paper for NAACL 2027 [PREPARING]: Evaluating how affective framing and emotional preambles modulate test-time reasoning compute and token duration across Qwen3.5-9B and Granite 4.2-8B, running extensive confound controls on mathematics and logic benchmarks for NAACL 2027.
  • Difftention [IN PROGRESS]: Investigating and implementing differential attention mechanisms to cancel background attention noise and sharpen focus on salient context tokens across long sequence horizons.
  • DKV + Attention [ACTIVE]: Fusing Differential KV-cache compression directly into the attention scoring operator to query low-rank key representations without standalone decompression roundtrips in memory.
  • Preparing DKV for MLSys (CUDA 4070 Super) [EVALUATING]: Running broader systems evaluation testbeds on an NVIDIA RTX 4070 Super (12GB) across extended context horizons to gather TTFT, decode throughput, DRAM memory footprint, and retrieval accuracy for MLSys.
  • Brain-Inspired Concurrent Cognition [EXPLORING]: Researching neuromorphic language model architectures where memory operates as an interconnected, distributed network rather than serialized KV stores, investigating asynchronous latent thought interaction and non-sequential convergence.

[02] Upcoming Horizons (2026–2027)

  • Research Collaborations & Lab Residencies: Actively exploring research residencies and systems engineering partnerships with frontier AI labs and systems groups on long-context efficiency, transformer memory systems, and custom kernel acceleration.
  • Visiting Research Fellowships & Academic Partnerships: Connecting with academic systems groups focused on efficient transformer inference and systems benchmarking.

[03] Field Dispatches & Engineering Observations

  • Why Anchor Tokens Prevent RoPE Phase Drift: In DKV, preserving an uncompressed initial anchor token per 256-token block anchors Rotary Positional Embeddings, preventing phase distortion while compressing remaining deltas.
  • Reference-Anchoring Isolates True Compression Loss: In CRBench, evaluating candidate context compression side-by-side against an identical dense reference insulates evaluation against intrinsic base model hallucinations.
  • Emotional Framing Expands Reasoning Tokens by +34.6%: Probing residual stream activations at Layer 18 reveals prompt anxiety triggers ungrounded verification loops without accuracy gain.
  • Judge Models Hide a 23.2% Reasoning Performance Spread: Conventional LLM-as-a-judge scoring compresses frontier model spreads due to agreement bias. Deterministic symbolic automata reveal authentic 23.2% spread.

Featured Systems & Engineering Projects

Flowject

Interactive, visual canvas-based workflow orchestration and automation engine. Built with robust real-time state synchronization, modular node graph evaluation, and execution scheduling.

Company Intelligence Engine

Real-time financial analytics platform analyzing market filings, news sentiment, and corporate entity knowledge graphs with automated summary extraction.

Mavis

Contextual desktop AI assistant integrating local language model execution with multi-modal vision-language inputs for offline developer productivity.

Signalist

Event-driven stock news signal pipeline pairing BM25 lexical keyword filtering with dense vector search to detect market-moving signals with low latency.

Honors, Awards & Competitions

  • Cognizance '25 (SnapSyntax) — IIT Roorkee: 1st Place / Zonals Winner (May 2025). Secured first place in technical algorithms and systems challenge at IIT Roorkee's annual technical festival.
  • Google — Gemini API Developer Contest: Contestant (May–August 2024). Built and deployed an asynchronous multi-modal AI application under live competition constraints.
  • Academic Merit & Research: B.Tech in Artificial Intelligence, Newton School of Technology (GPA: 8.0 / 10.0). Focus on GPU systems, high-performance computing, and efficient deep learning.

Technical Skills & Systems Expertise

LLM & Deep Learning Systems: PyTorch, Hugging Face Transformers, LLM Inference Optimization, KV-Cache Compression (DKV), Attention Kernels, Model Quantization, Low-Rank Truncated SVD, RAG Architectures, Mechanistic Interpretability, Representation Engineering.

GPU Computing & Performance Profiling: CUDA, Triton, Apple MLX, NVIDIA Nsight Systems, Nsight Compute, GPU VRAM Profiling, Memory Optimization, Kernel Benchmarking, Linux Systems Programming.

Programming Languages: Python, C++, CUDA C++, JavaScript, TypeScript, SQL, Bash.

Frameworks & Tools: FastAPI, React, Node.js, Git, Docker, Vector Databases (FAISS, ChromaDB), Linux CLI, Firebase.