I lead the design of distributed systems, query engines, and zero-to-one platforms — across Microsoft's Consumer, Commercial, and Hardware data universes.
"Consumer data is massive. Commercial data is deep. Hardware data isn't forgiving — and that's exactly where I operate."
Engineering lead and senior software engineer with 13+ years building distributed systems, query engines, and zero-to-one platforms — equally effective designing service architecture and communicating technical direction to executive audiences, up to the CEO.
I gravitate toward the problems where data is largest, least forgiving, and most consequential — consumer telemetry at petabyte scale, competitive intelligence reviewed by the CEO, hardware systems where a bad reading means a real-world failure. I believe the best platforms aren't built by choosing between depth and breadth — they're built by engineers who refuse to pick.
M.S. Digital Sciences, Kent State University · B.Tech. Electronics & Communication Engineering, JNTU
Drove a deep performance analysis of Fabric SQL against Trino — an engine that had never competed at this level — and the performance improvements that followed, directly shaping the engine roadmap. After my sign-off, executive reports migrated onto Fabric SQL, and the reusable benchmarks I built were adopted by multiple platform-engineering teams. The work was recognized all the way up to the Azure Data President and CEO Satya Nadella.
The lead engineer behind Microsoft's IDEAS journey onto Fabric — from the initial Trino-vs-Fabric evaluation to the semantic model strategy now powering analytics across 600+ teams. Designed the Direct Lake semantic model approach and the Fabric CI/CD lifecycle (Git integration, semantic workspace isolation, Azure DevOps promotion). Architecture published on Microsoft Learn as the enterprise reference. Automated promotion (−80% manual time), zero deployment regressions. Root-caused a platform-wide Power BI rendering issue — halved global first-render time (~15s → ~7s) for every report consumer across the globe.
Surfaced data exfiltration risk, then authored the end-to-end remediation — access-boundary design, asset-naming standards, and governance policies governing data flow and asset reachability. Presented the required API surface changes and requirements directly to the platform team; the design was adopted as a native Fabric capability, elevating an org-level safeguard into a product-wide security feature. Laid the foundational standards for EUDB-compliant data handling, defining how data residency and sovereignty controls are enforced at platform scale.
Architected the end-to-end competitive-intelligence platform powering the M365 Copilot analytics plane — automated ingestion pipelines, governed data models, and the executive analytics layer — spanning the GenAI landscape (OpenAI, Google, AWS, and others), security positioning, and cross-portfolio product analysis. Data-backed strategy with full automation from source to scorecard, delivering sub-five-second query performance on scalable, reliable data reviewed daily by executives, serving 600+ teams — with 100% metric parity across executive scorecards.
Designed and built a real-time platform for autonomous fleet management from scratch — a mission-lifecycle solution that turns operational alerts into autonomous action: the fleet navigates to the target, captures physical readings and imagery, and files them straight into the incident workflow, replacing manual on-site dispatch. Engineered for physical-world safety (layered prechecks, emergency-stop recall, automatic escalation) and hardened to key-less auth — cutting incident response to under ~10 minutes, architected to scale to N devices, and monitoring petabyte-scale telemetry (billions of rows daily) across global sites.
Scalable compute, reliable orchestration, and governed analytics — an end-to-end platform taking raw ingestion to governed, serving-ready data with lineage and data quality, on a two-cluster model.
A self-hosted enterprise analytics platform — a private data command centre where teams query live data, build dashboards, and use AI to accelerate analysis, with multi-provider auth configurable entirely from the UI.
Contributed use-case design and integration engineering for the official Trino .NET client library — an ADO.NET-compatible C# client open-sourced from Microsoft under the trinodb organization. Acknowledged at Trino Summit 2024.
A deep comparison of data platform performance across the Microsoft data ecosystem.
Process and strategy for versioning semantic models at scale.
Navigating semantic models and reports deployment challenges at scale.
Tips and tricks for semantic model optimization.
Performance optimization strategies for peak Direct Lake performance.
A long-awaited feature now at your fingertips.
Open to conversations on distributed data systems, query-engine architecture, and platform engineering.