Motorsport-to-Market Analytics -- EV Battery Research

Formula E is not a race series.
It is a live laboratory for battery R&D.

Julian Batto-Hokson Formula E Gen 1 to Gen 3, Lucid Motors, California EV registrations Python, SQLite, ChromaDB, Claude API, dbt 2026
Project Status: This is an active research and engineering project. The pipeline, database schema, and agentic query layer are built and functional. The analytical outputs described in this document are produced by the live system. Where proxy metrics substitute for unavailable telemetry, the substitution is documented explicitly.

Executive Summary

Formula E's core constraint -- a hard energy cap per race -- makes every race a controlled experiment in battery management. Teams that win by consuming less energy than rivals have demonstrated a transferable battery management insight. Teams that win by raw pace alone have demonstrated a setup advantage. This project separates those two signals.

Lucid Motors is the direct application case. As powertrain supplier to the Mahindra Formula E team, Lucid sits at a rare intersection: the same engineering team optimizing regen recovery and thermal management under race conditions also designed the Air Sapphire's 300 Wh/kg cell architecture. The FE2C framework quantifies that transfer across three generations of Formula E battery technology.

The central analytical question for California EV buyers: does battery density (larger pack, longer range buffer) or recovery efficiency (better regen, lighter pack) produce more real-world range improvement? Formula E's generational data -- where competitive pressure compressed development cycles that would take years in consumer contexts -- gives a data-driven answer.

Gen 1-3
Three battery generations analyzed: 28 kWh to Gen 3 regen-focused architecture
300 Wh/kg
Lucid cell energy density -- exceeds most competitors at market
250 kW
Gen 3 front axle regeneration -- the technical basis for the regen efficiency analysis
3 layers
Agentic RAG pipeline: SQLite star schema, ChromaDB vector store, Claude tool-use agent
8
Custom agent tools covering circuit analysis, efficiency scoring, CA market simulation
CA DMV
California registration data used to correlate racing innovation timelines with EV adoption

Why Formula E data can inform consumer EV decisions

01Every Formula E race is a controlled battery experiment

Unlike combustion motorsports where fuel consumption is an efficiency target, Formula E teams face hard energy caps that function like a fixed battery charge. Exceeding allocated energy means slowing down or disqualification. This constraint forces battery management optimization in a high-stakes environment where the cost of inefficiency is immediate and measurable.

GenerationBattery CapabilityConsumer EV Parallel
Gen 1 (2014-2017)~28 kWh -- teams swapped cars mid-raceEnergy density baseline -- proof of concept only
Gen 2 (2018-2020)~54 kWh -- full race on single charge achievedRange anxiety addressed for short-route EV use
Gen 3 (2023-present)Focus shifts to charge speed, lighter packs, active thermal managementAligns with consumer priorities: fast charging, range per kg

What this suggests for EV product teams

Gen 3's design philosophy shift from energy density to recovery efficiency is a direct signal about where battery R&D ROI is moving. The competitive pressure in Formula E compressed that shift into a single regulation cycle. Consumer EV teams tracking this transition have a leading indicator for where cell engineering investment should be directed in the 2025-2028 window.

02The density vs. recovery question has a California-specific answer

California presents two distinct EV stress environments that test different parts of the density vs. recovery question. The Inland Empire and Central Valley have high ambient temperatures (thermal management stress), stop-and-go traffic on I-10 and SR-99 (regen opportunity density), and sparse charging infrastructure (range buffer value). Bay Area and LA Metro have dense regen cycles but shorter average trip distances and higher charging frequency tolerance.

What the simulation produces

Dense urban areas (LA Metro, Bay Area) favor recovery efficiency -- more regen opportunities per mile and shorter trips where range buffer matters less than charging flexibility. Rural and exurban areas (Central Valley, Inland Empire) favor density -- fewer regen opportunities and larger distances between charging infrastructure make the range buffer more valuable. The FE2C simulation produces a ZIP-code-level recommendation map using CA DMV registration data weighted by these environmental factors.

Three-layer agentic pipeline

Layer 1 -- Structured Data

SQLite star schema

Fact table: race stints. Dimensions: battery hardware (Gen 1/2/3), track profile, manufacturer, environment. dbt layer enforces data quality: no negative lap times, energy proxies within declared capacity bounds, environmental completeness checks.

Layer 2 -- Semantic Search

ChromaDB vector store

Formula E technical documentation, FIA regulations, and Lucid engineering publications embedded and indexed. Allows the agent to answer questions that require qualitative technical context alongside quantitative race data.

Layer 3 -- Agent

Claude tool-use agent

8 custom tools covering circuit efficiency scoring, Delta-E calculation, regen opportunity indexing, cold-start circuit prediction, CA market simulation, and strategy recommendation by county. Answers multi-hop analytical questions against both layers simultaneously.

Architecture rationale

Formula E does not publish lap-level energy telemetry publicly. Proxy metrics are engineered from lap times, sector deltas, and gap data. The dbt quality layer documents every proxy assumption explicitly -- the same standard production analytics environments use when perfect data does not exist. Every metric in the analysis distinguishes between measured data and engineered proxy.

How efficiency is measured in this framework

M1Delta-E Score -- efficient battery management vs. costly position gains

Delta-E = Energy Consumed (proxy) vs. Positions Gained, calculated per stint using SQL window functions against the race average. A negative Delta-E indicates a team gained positions while consuming less energy than the stint average -- the signature of efficient battery management. A positive Delta-E indicates costly position gains that may be unsustainable over a full race distance.

What this means for consumer EV context

A negative Delta-E in a racing context maps to a real-world driver who gained range by driving more efficiently -- not by having a larger pack. Teams with consistently negative Delta-E scores have battery management strategies worth analyzing for consumer EV application.

M2FE2C Efficiency Score -- normalized cross-generation comparison

FE2C Efficiency Score = (Total Race Distance / Energy Used) x (Average Lap Velocity), normalized across generations using declared battery capacity differentials and adjusted for track profile using the aggression rating from the Track Profile dimension.

Modeling assumption -- documented

Normalization across Gen 2 and Gen 3 requires assumptions about how capacity differentials map to efficiency comparability. Those assumptions are documented in the dbt test layer and in the project write-up. Every efficiency comparison in this analysis specifies which generation it applies to and what normalization was applied.

Data sources and known gaps

SourceWhat it providesStatus
FIA / Formula E APILap times, race results, gap data, pit eventsLive -- ingested via pipeline
Open-Meteo APIAmbient temperature, humidity, track temperature at stint startLive -- weather enrichment layer
California DMVEV registrations by brand, powertrain, ZIP codeStatic -- loaded at pipeline init
Lucid Motors / FIA documentationBattery specs, Gen 3 regen specifications (250 kW front axle)ChromaDB -- semantic search layer
Formula E energy telemetryLap-level energy consumptionNOT AVAILABLE -- proxy engineered from lap time and gap data
Raw regen efficiency dataPer-stint regen recovery amountsNOT AVAILABLE -- Regen Opportunity Index constructed from deceleration zone analysis

Known data gaps and how they are handled

The two most important data gaps -- energy telemetry and regen efficiency -- are not available in any public Formula E data source. The proxy metrics that substitute for them are documented in the dbt quality layer with explicit assumptions. This is standard practice in production analytics environments where perfect data does not exist. The quality of a model is partly judged by how clearly its assumptions are stated.

LayerTool
Pipeline orchestrationPython -- 7-stage run_pipeline() function
DatabaseSQLite star schema (race_stints fact + 4 dimensions)
Data qualitydbt -- 10 targeted assertions, documented in test layer
Vector storeChromaDB -- Formula E and Lucid technical documentation
AgentClaude API tool-use -- 8 custom tools
CA market simulationPython -- ZIP-code weighted density vs. recovery recommendation

What this analysis cannot tell you

L Three things this project does not claim to prove

1. The efficiency scores are proxy-based, not measured. Formula E does not publish lap-level energy telemetry. Every efficiency metric in this project is derived from lap time deltas, gap data, and published battery capacity specifications. The proxy assumptions are documented in the dbt quality layer. Do not treat FE2C Efficiency Scores as equivalent to measured energy consumption data.

2. The Lucid-Mahindra technology transfer is assumed, not independently verified. The project uses published reporting on Lucid's role as Mahindra's powertrain supplier as the basis for the technology transfer thesis. The specific engineering details of what transferred, when, and how are not publicly available. The project treats this as a plausible case study, not a confirmed causal chain.

3. The California simulation is a weighted estimate, not a predictive model. The density vs. recovery recommendation by ZIP code is based on EV registration density and environmental proxies. It is a structured framework for thinking about the question, not a validated demand model. A/B testing or survey data from actual EV buyers in each region would be required to validate the recommendation.

Why these limitations are documented upfront

In production analytics environments, the quality of a model is partly judged by how clearly its assumptions and limits are stated. A stakeholder who understands what the analysis cannot tell them is better positioned to make good decisions than one who treats all outputs as equally certain. Every proxy metric in this project is labeled as such in the pipeline output.