Y3K Lab · Field Work

Solar Strategy

Production, storage, and a forecast graded daily.
Run a battery on a time-of-use rate so it makes money and lasts — then prove the whole thing live, in public.

an experiment in progress

connecting to the scoreboard…
Yesterday’s verdict
Tomorrow’s calls
The record

Electricity has a rush hour. At our house it starts at 4 PM — in summer every kilowatt-hour costs 2.24× what it did at noon — exactly when the roof’s solar is winding down. But at noon, the panels make more than the house can use. The whole experiment is one move: catch the cheap sun, hold it in a battery, and spend it during the expensive evening — automatically, every day, without wearing the battery out.

Part I — The Experiment

Run the battery on purpose

01

The question

Can a home battery save the most money and last the longest — automatically, every day, even in December?

The two halves fight. Saving money wants the battery full before every expensive evening; lasting wants it to spend as little time full as possible, because sitting at 100% is what ages its cells. A strategy is only interesting if it serves both at once — and keeps working on a rainy winter day. That tension is the experiment.

02

Three things you need to know

Everything below follows from three measured facts. First, the price of power follows the clock. We’re on PG&E’s EV2-A rate (NEM 2.0) — the hours never change, only the prices:

windowhourssummer (Jun–Sep)winter (Oct–May)
Off-peak00:00–15:00$0.25$0.25
Mid-peak15–16 + 21–24$0.45$0.41
Peak16:00–21:00$0.56$0.43
Fact 1 — the price spread is seasonal Summer peak is 2.24× off-peak — worth rearranging your life for. Winter peak is two cents above mid-peak — advising a family to shift their evening over $0.02 is theatre. Any signal you build should map on price, not on the tariff’s period names, or it cries wolf all winter.
Fact 2 — the sun is generous at the wrong time On a sunny day the roof makes more than the house can use at midday, and nothing by the evening peak. And under NEM 2.0, buy price = sell price in every window — an unnecessary off-peak grid charge is roughly free in dollars. Remember that for the failures below.
Fact 3 — batteries hate sitting full High state-of-charge is what ages NMC cells. That wear never shows up on a bill — it’s the invisible cost that shapes the battery-life half of every decision below.
03

Our idea — arrive full just in time

The hypothesis: charge from sunshine, arrive at 100% just in time for the expensive evening — never early — rest low overnight, and buy from the grid only as a cheap fallback when the sun falls short.

Stated as a spec, in priority order:

  1. A full battery before the expensive window opens, every day — the pack must carry the whole peak evening on stored energy (here: 100% by 3 PM, ahead of a 4 PM peak).
  2. Charge from solar wherever possible.
  3. Never charge from the grid at elevated rates — when the sun falls short, buy the gap only while power is cheapest (here: off-peak, before 3 PM).
  4. Minimum time at 100% — high state-of-charge is what ages NMC cells.
  5. Keep a reserve floor as outage insurance — some charge is never spent on economics (here: 20%, whole-home backup).

Goals 1 and 4 pull against each other, and that tension is the strategy. Here is what one working day looks like — measured on this system, where the evening peak consumes almost exactly one full pack (100% at 15:00 falls to the 20% reserve at ~21:20, minutes after peak ends):

overnight
Rest at the reserve (~20%). Not full. Resting low is the battery-life half of the design.
morning
Solar fills the pack. On a sunny day it reaches 100% with nothing else needed.
12:00
13:30
The fallback gates: if the sun is falling short, buy the gap from the grid while it is still cheap. Two gates because the charge curve is brutally non-linear — failure 2 below.
15:00 sharp
Stop all grid charging, drop the reserve to 20, and let the battery carry the house through the expensive evening. This step is the one that earns the money, and it must be unconditional and exactly on time — releasing early spends off-peak-priced stored energy that will be worth $0.56 in an hour.
04

What failed first (and what it taught us)

Our first automation ran flawlessly every night — and achieved none of its goals. Four findings, each paid for with real kilowatt-hours:

1

A backup reserve is a discharge floor, NOT a charge target.

The first automation set the reserve to 60% and enabled grid charging, believing the pack would stop there. It doesn’t — with grid charging on, a Powerwall charges to 100% regardless. The result: ~10.4 kWh bought nightly (twice the intent), 13.3 hours a day parked at 100% — identical to the behaviour it was written to eliminate. Trace your automations against measured state; a clean run log proves nothing.

2

The last 30% of the battery costs two-thirds of the charging time.

Measured: 20 → 71% runs flat at 5.7 kW and takes one hour; 71 → 100% tapers hard and takes two more. So “top up just before 3 PM” cannot work as a guarantee — a worst-case fill must start by noon. Hence two gates: 12:00 catches the deep deficit (SoC < 70%), 13:30 catches the shallow one (SoC < 95%) without committing early to power the afternoon sun was about to deliver anyway.

3

The costly failure was financially invisible.

Under NEM 2.0 the broken overnight charge cost roughly $0 — the midnight purchase and the displaced midday export were both $0.25. The real cost was battery wear: a full extra cycle plus 13 h/day at high SoC, the exact calendar-ageing the design existed to avoid. No dashboard tile will ever show that number. If you only watch dollars, this class of bug lives forever.

4

By noon, the battery already knows.

The elegant-sounding move is a solar forecast deciding the morning’s charge. But at 12:00 the battery’s own state of charge is the measurement of how much sun actually arrived — no forecast needed for the fallback decision. The forecast earns its keep elsewhere (Part II), and only if it proves it can beat this simple rule.

05

Does it survive winter?

A strategy that only works in July isn’t a strategy. From 1,041 days of production records (Part II explains how we got them):

Winter fills the pack — thinly December median: 17–18 kWh/day against a 10.4 kWh pack. Winter solar usually fills the battery. The fallback gates fire more often; that’s the design working, not failing.
Genuinely rainy days are a buy The ten worst days delivered 2–4 kWh before 3 PM. On those days solar cannot come close, and buying the whole gap at off-peak is simply the correct move.
Summer evenings are the whole game On this rate, one badly-timed hot evening imported $0.51 of peak power; normal days import $0.00–0.02. The 15:00 release, on time, is where the money is.
Also in the room A VPP guard — during a Virtual Power Plant event the utility pays the battery to discharge, so automations stand down (phrased as not-active, because the sensor reads unknown until the first event ever fires). The 20% reserve is untested outage insurance by design. And NEM 3.0 changes the arithmetic entirely — export credits collapse, and the forecast-driven version becomes the money-maker.
Part II — The Second Experiment

We Fit Our Own Sky

The rules above are simple on purpose. Could a weather forecast run the day better? Maybe — but first it has to prove it, in public.

06

The second question

Can we predict tomorrow’s solar, for this exact roof, better than a commercial forecasting service?

Our approach: replace a textbook solar-forecast formula with a curve fitted to our own roof, our own weather station, and 2.8 years of our own production history. The model runs twice a day on a Mac mini, predicts tomorrow’s solar in kWh, writes every prediction to a database beside what actually happened, and scores itself. A commercial forecast (Forecast.Solar) runs beside it on the same leaderboard, as a rival — never a target.

07

How we built the predictor

Forecasting solar production sounds like one hard problem, but it decomposes into three arrows — and two of them were already solved on this homelab:

forecast cloud % → forecast sunlight (GHI)the ONLY empirical unknown forecast GHI → irradiance on OUR roof deterministic astronomy roof irradiance → AC kilowatt-hours measured system efficiency

The astronomy (NOAA solar position → Haurwitz clear-sky → Erbs diffuse split → Hay–Davies transposition onto the real 23° / 245° + 155° roof planes) already ran at ingest, validated to 0.0000% against an identity control.

The efficiency is measured end to end — actual kWh ÷ clear-model kWh over recent verified-clear days — one number that absorbs every systematic bias in the chain and re-calibrates itself as seasons change.

The transmission arrow — how much light this valley’s clouds actually block — is the only thing that had to be learned, and it is exactly the thing a generic vendor model cannot know about one microclimate. Hourly cloud forecasts come free from two providers already in Home Assistant (met.no and OpenWeatherMap, 48 h ahead), which also makes them competing inputs, scored against each other.

The shortcut — 2.8 years of data we already owned. The Tesla app renders production history back to install day, and the Fleet API serves the same data: 1,165 days (2023-05-30 → present), pulled in one run, verified three independent ways — the lifetime total agrees with the documented lifetime counter to 0.1%; overlap days match our own meter to a consistent ~3%; and the record starts on exactly the install date, with empty months before it. Each day was labelled actual ÷ clear-sky ceiling — its transmission ratio — and paired with ERA5’s hourly observed cloud for the same dates. That’s the training set that would otherwise have taken months of waiting for cloudy days to accumulate. One boundary was measured, not assumed: the first four months ran export-capped pending PG&E permission-to-operate, so those 124 days were deleted — they measure the cap, not the sky.

08

The results — our sky beats the textbook

The fitted transmission curve omfit_v1
T = 1 − 0.36 · (c/100)5.20
how much light this sky passes at cloud cover c
modelholdout MAE (all 218 days of 2026, never seen during fitting)
Fitted: T = 1 − 0.36·(c/100)5.203.97 kWh/day
Kasten–Czeplak textbook seed5.14
Cloud-blind (ignore clouds entirely)5.28

23% better than the textbook, 25% better than ignoring clouds — on data the fit never touched. Training was 823 post-PTO days paired with hourly observed cloud from ERA5 reanalysis, efficiency always solved from the train set so the holdout stays honest.

a = 0.36 is the genuinely interesting part: the textbook curve (fit to 1980s Hamburg) assumes clouds block 75% of light at full cover — our sky, as ERA5 sees it, passes roughly twice what Hamburg’s did. Marine-layer optics, plus grid-cell “overcast” rarely meaning locally black. That one number is why the seed was systematically wrong here, and it’s knowledge no vendor model would have handed you.

And the first real test was scheduled by the weather: the first cloudy forecast day split the sources for the first time (met.no 29.2 kWh vs OWM 24.6) — until then every prediction had been a sunny-day handshake. One caveat stays standing: the curve was trained on reanalysis cloud and runs on forecast cloud — that gap is exactly what the daily scoring loop exists to measure.

Transmission — cloud cover in, light out fitted to our sky textbook, at full cover
At 100% cloud the textbook passes 25% of clear-sky light; our fitted curve passes 64% — the marine-layer difference the seed could never have known. Scored days will accumulate against this curve as the record grows.
09

How we keep ourselves honest

The rules the pipeline enforces on itself — the part worth publishing. They stop being claims the moment the scoreboard below goes stale-checkable.

Score against reality, never against a peer.

Forecast.Solar runs on the same leaderboard, configured honestly with the real roof geometry — but only actuals ever tune our model. Tuning to a rival teaches imitation, not prediction.

Hold out a year.

The curve was fitted on 2023–2025 and judged on 2026 only. The efficiency scale was always solved from the train set — the holdout stayed untouched until the final comparison.

Version every prediction.

Rows are tagged with the model that made them (kc_seed_v1, omfit_v1, vendor), so no average ever mixes eras. When the model is refit, this page shows the era boundary — never a merged score.

Predict, record, and only then act.

The forecast controls nothing. The battery-charging automation keeps its simple, static rules until the projection demonstrably beats them on decision accuracy — “would we have needed the grid?” — not just kWh error. House rule: a new job reports; it does not act. Authority is earned.

Say “not enough data” out loud.

Until the record contains genuinely cloudy days, accuracy numbers are vacuously good — sunny-day agreement tests the astronomy, not the forecast. The daily report states this itself, and so does the scoreboard below.

10

The scoreboard — live, graded twice daily

This is the results section that never stops updating. Every run writes predictions for today and tomorrow from three sources — our model on met.no clouds, our model on OWM clouds, Forecast.Solar — then back-fills what actually happened onto every completed day and reports MAE per source and per horizon. Day one of the benchmark: for a forecast-sunny tomorrow, our model said 39.3 kWh, Forecast.Solar said 27.7 — against recent sunny-day actuals of 39–41. The leaderboard settles it from here.

Leaderboard — mean absolute error, kWh/day
connecting…
Predicted vs delivered — the last 30 days the sun delivered ours · met.no ours · OWM Forecast.Solar
connecting…
Ticks are the recorded calls, day-ahead where the horizon is tagged; Forecast.Solar appears at day-ahead only — its evening same-day values are hindsight-contaminated. Bars are what the roof produced. Days without a bar aren’t scored yet — predictions stand recorded, waiting for their sun.
Today, live — solar & battery solar W battery %
connecting…
Solar and battery state only, by design. Live household load is an occupancy signal, so home and grid appear on this page only as daily totals — a privacy constraint, not a gap.
11

What happens next

The forecast controls nothing today — that’s rule 04, not an accident. The battery keeps its simple, static rules until the projection demonstrably beats them on decision accuracy — “would we have needed the grid?” — across a window that includes genuinely cloudy days. If it wins, it earns the noon decision, with the static rule kept as the fallback. Either way, the scoreboard above keeps score in public.