Skip to content

Methodology

How Native Refuge builds your report

Every Native Refuge report is selected for one specific address — your county, your ecoregion, your soil, and your climate, ranked by the wildlife those conditions support. (Your watershed and hardiness zone frame the report as context; they don’t filter the plant list.) The pipeline combines federal ecological data, peer-reviewed wildlife frameworks, and AI-curated native-range data cross-checked by multiple independent models. Here’s how, step by step.

What makes Native Refuge different

Property-level, not ZIP-level.

Most free native-plant lookups (Audubon’s, the National Wildlife Federation’s, Homegrown National Park’s) return one list per ZIP code or per state region. Two homes a mile apart get the same recommendations even when one sits on a valley-floor alluvial loam and the other on a serpentine ridge top with completely different native plant communities.

Native Refuge resolves your address to one specific point and runs an EPA Level IV ecoregion lookup, a county-level native-range filter, an elevation match against each species’ published range, a soil-substrate filter (no serpentine specialists on alluvial soils, no coastal-dune species inland), and a 30-year PRISM climate match against each species’ own climate niche. The plant list you read is the intersection of those filters — not an averaged regional list. Your watershed and the rest of your property data are drawn alongside as context; they frame the report rather than filter it.

Every species in your report cleared those filters. Every species filtered out was filtered for a stated, traceable reason — not buried in a one-size-fits-all list.

Step 1

Address → place

Your address is geocoded to a single point, then resolved against three datasets in parallel:

  • County, via PostGIS point-in-polygon against TIGER/Line 2024 boundaries.
  • EPA Level IV ecoregion — one of about 180 fine-grained ecological units that subdivide California by climate, geology, soils, and historical vegetation. A Level IV ecoregion is finer than a USDA hardiness zone: two yards 30 miles apart can sit in entirely different ecoregions, and the native plant communities of each are different.
  • Property context — elevation (USGS), soil texture, series and drainage (USDA SSURGO), 30-year climate normals (PRISM 4km), surface-water proximity (USGS NHD), watershed (USGS Watershed Boundary Dataset HUC12, walked up to sub-region and terminal water body), and USDA hardiness zone (PRISM 2023).

Step 2

Building the candidate pool

Before any plant is considered for your report, it must pass two native-range gates and two practical filters.

  • State-level gate (NatureServe). Where NatureServe Explorer has assessed the species, it must list California among its native states; species it hasn’t assessed aren’t blocked here — the stricter county gate below is the real test. This catches obvious wrong-state errors.
  • County-level gate (AI-curated, multi-vendor verified — Step 3 below). The species must be documented as genuinely native to your specific county, not just somewhere in California.
  • Tier: keystone or supporting species only. Tier 3 ornamentals with limited wildlife contribution are excluded.
  • Commercial availability: only plants stocked by California native-plant nurseries, reconciled against Calscape’s current carrying-nursery data. Recommending a plant a customer can’t source doesn’t help.

Step 3

County-level native-range curation

The county-level gate is the part of the pipeline we invested the most in. A species being “California-native” is not enough — California spans coastal redwood forest, alpine meadow, oak savanna, chaparral, and four deserts. Recommending a Sierra Nevada montane species for a downtown Sacramento yard is the kind of error that makes a report feel auto-generated.

So we built a county-by-county native-range dataset for the entire recommendable pool — every species, all 58 California counties — rather than relying on the state-level lists that are freely available. It is assembled with multiple independent AI passes, cross-checked against each other, and grounded in the standard botanical references (Jepson eFlora, CNPS Calscape, BONAP, USDA PLANTS). Independent cross-verification is the point: agreement between passes built on different foundations catches errors any single pass would repeat.

Every cell carries explicit provenance — what asserted it, with what confidence, and a habitat note where one applies. No hidden manual lists; every choice is traceable, and the record is what lets us keep auditing and correcting it. That correction work is continuous: the dataset has been through repeated botanist-reviewed passes that removed cells the evidence didn’t support.

What that buys you: redwoods don’t appear in Sacramento Valley reports, giant sequoias don’t appear at coastal addresses, and coast live oak doesn’t appear in inland desert reports. We hold the list to that standard and keep testing it against known-native benchmarks — but the honest framing is “rigorously filtered and actively corrected,” not “provably perfect.” If a plant in your report looks wrong for your site, tell us: there’s a link on every plant, and those reports feed the same correction queue.

Step 4

Address-specific filtering

The candidate pool from Steps 1–3 is then filtered against your specific address:

  • Elevation — species whose elevation range doesn’t overlap your property are excluded, with a 300-foot buffer in each direction. Species that the per-species curator vouched for your specific county get an extended 500-foot buffer — letting foothill-fringe species appear for sea-level addresses near their range edge.
  • Substrate exclusion — species whose plant communities don’t fit your soil are excluded. Serpentine endemics are excluded on non-serpentine soil; wetland obligates are excluded on upland soil; coastal-strand species are excluded inland.
  • Synonym suppression — known taxonomic synonyms are collapsed to their canonical species so a report doesn’t show two cards for the same plant under different botanical names.
  • Climate match — most species carry a climate niche: the precipitation and temperature range of their California occurrence records, from PRISM 4km normals. A species whose niche doesn’t overlap your property’s 30-year climate is excluded. This is the desert / montane discriminator — it’s why a Palm Springs report leads with mesquite and palo verde, not coastal oaks.
  • Local corroboration — a species with no climate niche on file must instead be backed by verified GBIF occurrence in your ecoregion. A county-native still surfaces where observation density is thin, but a name with neither a matching climate niche nor local records is held back rather than shown on faith.

Step 5

Wildlife-value ranking

Within the eligible pool, species are ranked by a composite wildlife score that combines four taxonomically distinct signals, so a plant’s rank reflects its impact on insects, native bees, broader pollinators, and birds — not just one of those.

The four signals, in order of how much they count:

  • Caterpillars — the largest single factor. California-specific counts of moth and butterfly species whose larvae feed on this plant, sourced from CNPS Calscape where available, falling back to Doug Tallamy’s national genus-level data for the ~15% of species without a Calscape entry. The keystone-genus framing (oaks, willows, cherries, asters, goldenrods support an outsized share of the insect biomass that feeds nesting songbirds) anchors the whole approach.
  • Specialist native bees — second. Per-genus count of California-distributed oligolectic bees (bees that depend on a single plant genus or family for pollen) from Jarrod Fowler’s 2020 compilation. Highest for late-season composites — goldenrods, gumplant, asters — and zero for wind-pollinated genera like oaks and pines.
  • Pollinator breadth — third. Per-species count of distinct flower visitors (bees, hoverflies, beetles, butterflies, and more) from GloBI (the Global Biotic Interactions network). Captures pollinator diversity beyond specialists, log-normalized to blunt the sampling bias toward common showy plants.
  • Bird food — the smallest of the four, because it’s a yes/no trait rather than a count. A hand-curated tag (berries, acorns, nuts, seeds, samaras) for the California natives known to provision birds directly. There is no defensible public dataset for species-level CA bird-plant relationships, so we curate rather than fake it.

Two small nudges on top:

  • Keystone species get a soft preference inside their layer, so the highest-impact plants surface first.
  • Range confidence — species documented as solidly within their native range in your county rank slightly above ones at the edge of it.

Each plant card surfaces the underlying signals as wildlife icons — a caterpillar, bee, bird, and hummingbird mark — so you can see at a glance why a plant ranks where it does. The exact caterpillar-host and specialist-bee counts live on each plant’s detail view. The composite score sets the default order; the icons, with that full detail, are the audit trail.

Step 6

Layered presentation

Species are grouped into five structural layers so the list reads as a complete planting palette, not a wall of one growth form: trees, shrubs, perennials & wildflowers, ground covers, and native grasses.

Within each layer, a genus cap keeps the list diverse: three species per genus for keystone genera (oaks, willows, manzanitas, ceanothus, pines, currants, salvias, penstemons, eriogonums, lupines, monkeyflowers — the regional anchors most customers expect to see well-represented), two per genus elsewhere. Inside a genus the cap picks the species that’s actually most abundant in your ecoregion, not the one with the highest state-level wildlife score — so a Sacramento Valley address gets Valley Oak, not the foothill-dwelling Blue Oak with a near-identical score but a fraction of the local presence.

Final ordering within the report: layer → composite wildlife score → commercial availability (widely-stocked plants edge out specialty-only ones at the same wildlife rank) → botanical name.

Step 7

The property-specific context

The plant list is the core of the report, but Native Refuge also renders four pieces of context that frame why these species and not others:

  • Aerial view of your land — a satellite tile centered on your geocoded coordinates so the report begins with the actual place, not an abstraction.
  • Property data panel — elevation, soil texture and series, USDA hardiness zone, 30-year rainfall and temperature, and the watershed your rain drains into. Each field hides cleanly when its source data is unavailable.
  • Watershed cascade — your USGS HUC12 sub-basin, walked up to the sub-region and terminal water body (Pacific, San Francisco Bay, Sacramento Delta, Salton Sea, etc.). Native plantings’ deep roots help rain soak in rather than run off the way it does over shallow-rooted lawn, keeping pollutants out of the rivers below.
  • Nearest native nurseries — the closest California native-plant retailers within about 50 miles, drawn from a curated directory of CNPS-listed nurseries. Coverage is uneven across the state; where we have none in range, the report links the CNPS statewide directory instead.

Sources

Data provenance

Every Native Refuge recommendation traces back to public data. The full source list:

  • EPA Level IV Ecoregions of California — United States Environmental Protection Agency via USGS.
  • TIGER/Line 2024 County Boundaries — U.S. Census Bureau; used for the point-in-polygon county lookup.
  • NatureServe Explorer — state-level native-presence records; the Layer-0 native-range gate.
  • County-level native-range curation — AI-curated dataset using multi-model verification, built on training-data knowledge of Jepson eFlora, CNPS Calscape, BONAP, and USDA PLANTS.
  • GBIF occurrence records — Global Biodiversity Information Facility, research-grade observations.
  • CNPS Calscape — per-plant pages; primary California-specific lepidoptera host count for most of the recommendable pool (the rest fall back to the national Tallamy data below).
  • Tallamy & Shropshire 2009/2020 — national genus-level lepidoptera host data via the National Wildlife Federation Native Plant Finder; the lep fallback for species without a California-specific count.
  • Jarrod Fowler 2020 — Pollen Specialist Bees of the Western United States; the specialist-bee signal.
  • GloBI (Global Biotic Interactions) — flowersVisitedBy records; the pollinator-breadth signal. Open data under CC0.
  • Bird-food categorization — list of ~145 California natives with documented direct bird-foraging value — a hand-curated core (CNPS, Audubon, and ornithology references) extended by ecological-role inference.
  • Calflora and Jepson eFlora — native status, growth form, horticultural traits.
  • USDA NRCS SSURGO — soil texture, series and drainage.
  • USGS Elevation Point Query Service — per-address elevation.
  • PRISM Climate Group — 30-year climate normals (4km grid).
  • USGS National Hydrography Dataset — surface-water proximity.
  • USGS Watershed Boundary Dataset (HUC12) — sub-basin lookup and the parent-region cascade drawn at the top of every report.
  • PRISM Climate Group + USDA — 2023 USDA Plant Hardiness Zone Map (800m grid).
  • NOAA NCEI 1991-2020 Climate Normals — per-station median spring and fall freeze dates, growing-season length.
  • CNPS native plant nursery directory — the curated list of California native-plant retailers feeding the "Where to buy" section.
  • Mapbox Satellite — aerial imagery for the report’s opening view, attribution displayed beneath the image.
  • iNaturalist and Wikimedia Commons — research-grade photographs and per-photo attribution, used under their Creative Commons licenses with photographer credit on each image.

iNaturalist and Wikimedia Commons images are used under their Creative Commons licenses (CC0, CC-BY, CC-BY-SA, public domain), displayed unmodified with photographer credit. Embedding within Native Refuge reports is not a derivative work.

Limits

What Native Refuge does not do

Native Refuge is a vetted shortlist, not a landscape design. Placement, hardscape, shade, succession, and cultivar selection are all out of scope — where the wild form differs substantially from common trade cultivars, the plant’s detail view notes it. The list starts you right — your eyes on the ground finish it.

Native Refuge is not a substitute for professional landscape design, arboricultural consultation, or fire-defensible-space planning.

V1 is calibrated for Mediterranean and Valley California — the climate zones most California gardens occupy. Desert and high-Sierra addresses are supported too: a PRISM climate filter keeps each report climate-appropriate (a Palm Springs address leads with mesquite, palo verde, and ironwood — not coastal or montane trees). Where a site is hyperarid enough that no cultivable native qualifies, we return an explicit limited-coverage notice rather than a forced list; where the list is real but unusually short, we say so on the page before you pay, with the count and the reason. These climate types saw less validation than the Mediterranean core, so we’re continuing to refine them post-launch with signal from real customer reports.