Features Export Specification
Status: Implemented (Living Spec)
Owner: WEPPpy NoDb export subsystem
Primary module: wepppy/nodb/mods/features_export
Replaces legacy export modules: wepppy/export/gpkg_export.py, wepppy/export/prep_details.py, and associated route/task wiring
Document posture: Living working specification; mutate when implementation evidence shows a better contract or exposes gaps.
1. Summary
Create a NoDb features_export mod for user-configurable spatial and spatial-temporal exports across WEPP, Omni, Ash/WATAR, WEPP interchange, SWAT interchange, and AgFields datasets.
This is an immediate replacement for legacy gpkg/gdb export behavior, but implemented with NoDb controller patterns, canonical RQ polling contracts, and dependency-aware cache reuse.
AgFields support is parity+ (spatial + WEPP interchange metrics), including automatic on-demand AgFields interchange generation when required for requested export layers.
Data extraction and merge orchestration for export payload assembly is DuckDB-first (SQL joins/projections/filters) for performance and deterministic schema control; pandas merge loops are non-compliant for production payload assembly paths.
The normative materialization architecture is key-first and geometry-last: build one attribute table per carrier/context/scope keyed by canonical ids, then attach geometry exactly once from canonical carrier geometry.
This architecture is the default implementation contract (no temporary feature-flagged parallel path).
User-facing dataset labels and output layer names must prioritize established WEPP output vocabulary over internal family or implementation tokens.
2. Supported Formats
geojson(single-layer format)geoparquet(single-layer format)parquet(single-layer geometryless tabular format)csv(single-layer geometryless tabular format)kmz(single-layer format)geopackage(multi-layer container format)geodatabase(multi-layer FileGDB container via GDALOpenFileGDB)
Format token contract:
- Canonical request token is
geodatabase. - Backward-compatible alias
f_esriis accepted and normalized togeodatabase. - FileGDB payload member extension remains
.gdb.zipinside the final download bundle. - Geodatabase creation explicitly selects GDAL's built-in
OpenFileGDBdriver;FileGDBis not an accepted creation-driver alias because GDAL 3.10 still resolves it to the optional Esri SDK driver. - OpenFileGDB uses its default broad ArcGIS compatibility mode. Integer64 source fields may therefore be represented as Float64 in the FileGDB output; the writer does not set the ArcGIS Pro 3.2+ compatibility option.
Packaging rules:
- All format downloads are
.zipartifacts. - Single-layer formats produce one file per resolved layer inside the zip bundle.
- Multi-layer formats produce one container payload member inside the zip bundle.
- KMZ is single-layer only; multi-layer requests produce multiple
.kmzfiles in the zip. - Geometryless formats (
parquet,csv) always emit tabular outputs without geometry columns/encodings. - For geometryless formats, required identity/join fields remain included even when geometry is removed.
- Geometryless formats support optional
tabularcontrols:concatenate_tables=trueconcatenates hillslope carrier tables into onehillslopesfile and channel carrier tables into onechannelsfile.temporal_layout=wide|longcontrols temporal measure shaping forevent/yearlylayers (widedefault).
- Geometryless writer path is table-native end-to-end: tabular exports consume DataFrame payloads directly and must not serialize/parse FeatureCollection JSON in the writer path.
- Geometryless carrier materialization is independent of geometry files: tabular outputs are produced from attribute sources only and do not enrich identity columns from carrier geometry datasets.
- Every zip artifact must include:
- export payload members (data files/container members),
manifest.json,- generated
README.md(artifact metadata summary).
- Artifact bundles must not include
profile.ymlor built-in profile files; profile discovery/replay is route-level (profile/resolve) and publication-level (published/index.json) metadata. - Identity normalization contract for all output formats (geometry and tabular):
- Emit canonical identity columns
topaz_id,wepp_idas the first two output columns. - Coalesce identity aliases (
TopazID/topaz_id,WeppID/wepp_id) into canonical columns. - Remove redundant alias columns after coalescing.
- Emit canonical identity columns
3. Layer Catalog And Discoverability
The backend owns an authoritative, versioned export layer catalog. GL Dashboard taxonomy and discoverability are used as a blueprint, but export contracts are backend-defined and independent of frontend module internals.
Catalog source of truth:
- Machine-readable catalog file:
wepppy/nodb/mods/features_export/layer_catalog.yaml. - Catalog header block lives under
metadataand includes versioning/compatibility fields. - Runtime layer discovery for export/UI must read from
layer_catalog.yaml, not hardcoded layer maps.
Catalog contract:
- Top-level keys:
metadata: catalog metadata/version header block.layers: array of layer definitions.metadataminimum fields:catalog_versionschema_versionupdated_at_utcownerstatus(draft|active|deprecated)resolver_contract(allowed locator kinds and template-variable contract)- Layer definition minimum fields:
layer_idfamilyscope_class(scope_aware|scope_invariant)geometrycontract (type, locator, feature-id metadata; geometry locator participates in readiness and dependency fingerprinting)joincontract (primary_key, optionalfallback_keys, optionalsource_key_map)sourcesdatapaths (direct data sources used to build the layer)dependencies(additional files that must participate in readiness/fingerprint checks)temporaloptions (supported_modes,grain,time_columns,mode_rules)measures.requiredandmeasures.optionalcolumnscontract for UI-visible field selection:column_id(canonical source/output field key)label(human-readable field name)unitmetadata (display_unit,unit_class,is_unitized)default_selected(boolean)- optional availability/selector guards
- When a layer omits an explicit
columnsblock, runtime schema discovery (from resolved source datasets) is the fallback source of truth for UI column selectors and unit labels.
- SWAT table-profile contract for
swat.interchange.*layers (table_profiles, profile-level geometry/join rules, non-spatial behavior) - Optional measure availability rules (
requires_any_column,requires_all_columns, version gates, selector constraints) - Version-gate semantics:
min_source_versionis a semantic version string compared against the resolved dependency manifestversionfield (for exampleinterchange_version.json.version). - Missing optional measure behavior (
warnwithmeasure_unavailable)
Locator contract (strict):
- Every
geometry.locator,sources[*].locator, anddependencies[*].locatoruses exactly: kind: one ofnodb_ref|relpath|path_templatevalue: locator value string- Locator aliases such as
path,path_ref,path_template, orsource_refare not allowed. path_templateexpansion variables are defined inmetadata.resolver_contract.path_template_vars.- Resolved locator paths must stay inside allowed dependency roots:
- default allowed root is the active working directory (
wd); - for canonical Omni child runs (
_pups/omni/scenarios/*and_pups/omni/contrasts/*), the parent run root (path segment before_pups) is also allowed.
- default allowed root is the active working directory (
- Canonical dependency
relpathvalues are recorded relative towd; parent-run references are therefore expected to include../segments for Omni child runs.
Initial layer families:
- Watershed: subcatchments, channels.
- Landuse: dominant class and coverage attributes.
- Soils: dominant class and physical properties.
- Ash/WATAR: hillslope ash transport outputs.
- AgFields spatial: field boundaries and sub-field polygons.
- AgFields WEPP metrics: sub-field/field metrics sourced from
wepp/ag_fields/output/interchange/*. - WEPP: canonical output datasets labeled with familiar file names (for example
H.element.parquet,H.wat.parquet,H.pass.parquet,H.loss.parquet,H.soil.parquet,return_period_events.parquet,chan.out.parquet,chanwb.parquet). - SWAT interchange:
swat/outputs/run_*/interchange/*. - Omni scenarios:
_pups/omni/scenarios/*. - Omni contrasts:
_pups/omni/contrasts/*.
Discoverability requirements:
- Internal ids like
wepp.temporal.eventsare backend-only tokens and must never be shown as the primary UI label. - Primary dataset labels are catalog-owned via per-layer
labelinlayer_catalog.yaml; route/controller code must not maintain a parallel hardcoded label map. - Group rendering is discovery-driven: hide groups with zero currently available datasets for the active run/config.
- Group order keeps Omni families at the bottom of the catalog list.
4. Output Scope Contract Alignment
output_scopes is an array selector with values baseline|roads.
Default is ["baseline"].
Scope resolution must follow the canonical output-scope contract:
- Only datasets rooted at
wepp/outputare scope-rewritten forroads. - Paths outside
wepp/outputremain unchanged for both scopes. - Scope values are normalized case-insensitively to canonical lowercase.
- Duplicate scopes are deduplicated before execution and cache-key hashing.
- Invalid scope values fail with 400; silent fallback is forbidden.
- Catalog
scope_roottoken mapping is fixed tobaseline=outputandroads=roads/output.
Layer scope classes:
- Scope-aware layers: WEPP summary, WEPP temporal, WEPP interchange layers rooted at
wepp/output. - Scope-invariant layers: watershed, landuse, soils, ash/watar, AgFields, Omni, SWAT interchange, and any layer not rooted at
wepp/output.
Export behavior:
- WEPP and Omni output datasets are consolidated by geometry carrier per context to keep layer counts legible:
- Up to one
sbs_map-subcatchmentslayer per scope. - Up to one
chan_map-channelslayer per scope.
- Up to one
- Base WEPP context emits at most two consolidated layers per requested scope (
subcatchmentsand/orchannels) depending on selected outputs. - Each selected Omni scenario/contrast emits its own consolidated
subcatchmentsand/orchannelslayer set per requested scope. - Scope-invariant families remain single-emission artifacts and are not duplicated per scope unless explicitly scope-aware in catalog metadata.
- Consolidated layer names must use descriptive run/context naming:
- Baseline base context:
{runid}-sbs_map-subcatchments,{runid}-chan_map-channels. - Roads scope:
{runid}-roads-sbs_map-subcatchments,{runid}-roads-chan_map-channels. - Scenario context:
{runid}-scenario-{scenario_id}-{carrier}. - Contrast context:
{runid}-contrast-{contrast_id}-{carrier}.
- Baseline base context:
- If one requested scope is missing for a scope-aware layer, export available scopes and emit warning code
scope_missing_layer. - If a requested scope is not applicable to a scope-invariant layer, emit warning code
scope_not_applicable. - If no layer resolves after scope processing, return 404.
- When
format=csv|parquetandtabular.concatenate_tables=true, carrier-concatenated rows include provenance columns:output_scopeomni_scenario(scenario-context rows only; null otherwise)omni_contrast_id(contrast-context rows only; null otherwise)
tabular.temporal_layoutbehavior forevent/yearlytemporal modes:wide(default): one row per feature key with temporal selector tokens appended to measure column names.long: one temporal selector column (date,return_period, oryear) with measure columns across multiple rows.
Carrier materialization contract (normative):
- Consolidated carriers are built in two phases:
- Phase A (
DuckDB attribute core): materialize one table per{context, selector_id, scope, carrier}from discovered datasets using canonical join keys. - Phase B (
geometry attach): join Phase A output to a canonical carrier geometry table exactly once.
- Phase A (
- Canonical key precedence:
- Subcatchments carrier:
topaz_idpreferred,wepp_idfallback. - Channels carrier:
chn_idpreferred,topaz_idfallback. - Catalog
join.source_key_mapoverrides remain authoritative for source-specific key resolution.
- Subcatchments carrier:
- Contract clarification:
wepp.summary.channelsmust materialize its internal metrics-plus-attributes join onwepp_idbecauseloss_pw0.chn.parquetis keyed by WEPP channel id while the canonical channel geometry carrier remains topaz-facing.
- Each source dataset must be reduced to one row per effective carrier key before joining into Phase A output (via deterministic temporal filtering, deterministic projection, and deterministic dedupe/aggregation rules when needed).
- Unresolved many-to-many key joins on a carrier hot path are contract violations and must fail explicitly with
materialization_error; silent Cartesian growth is forbidden. - Legacy non-carrier source merges must resolve identity keys from explicit join contract candidates (
join.primary_key,join.fallback_keys,geometry.feature_id_keys). If no candidate resolves, fail withmaterialization_error; arbitrary first-column fallback is non-compliant. - Canonical carrier geometry tables must contain one geometry row per effective carrier key. When raw geometry sources contain repeated key rows, geometry must be canonicalized (for example, deterministic dissolve/aggregation) before Phase B.
- For spatial carriers, the canonical geometry keyset is the authoritative export row domain. Phase A keys not present in canonical geometry are excluded before final attachment, and final row/feature counts must match canonical carrier entity counts.
- Repeated geometry-attached frame merges (geometry-first per-dataset pipelines) are non-compliant for production export paths.
5. API Contract (rq-engine)
5.1 Submit Export Job
POST /api/runs/{runid}/{config}/export/features
Auth contract:
- Submit/profile-resolve endpoints require
rq:exportplus run-access authorization. - Download endpoints (
job/{job_id}/download,published/{profile}/download) allow anonymous access for public runs; non-public runs requirerq:exportplus run-access authorization. - Polling auth follows canonical
/rq-engine/api/jobstatusand/rq-engine/api/jobinforoute policy;features_exportdoes not introduce route-specific polling auth overrides.
Transport contract:
- Request body is
application/jsononly. - Unsupported content type (including
multipart/form-data) is rejected with 415. - Missing JSON body, empty JSON body, or query-only submissions are rejected with 400 validation errors.
- This avoids
FormDatalist-collapsing ambiguity in shared payload parsing.
Request schema:
format: required enum from Section 2.units: required enumsi|english|project.crs: optional enumwgs|utm, defaultwgs.layers: required non-empty array of layer IDs.output_scopes: optional non-empty array ofbaseline|roads.scenarios: optional non-empty array of Omni scenario IDs.contrast_ids: optional non-empty array of Omni contrast IDs.- Backward-compatible aliases
scenarioandcontrast_idare accepted and normalized into single-entry arrays. swat_run_id: optional SWAT run selector, defaultlatest.swat_tables: optional object with one ofincludeorexclude, each an array of table names.column_selection: optional object keyed bylayer_idwith one of:include: non-empty array ofcolumn_idvalues to export.exclude: array ofcolumn_idvalues to drop from export.includeandexcludeare mutually exclusive per layer.
temporal: optional object.temporal.mode: optional default enumannual_average|yearly|eventused when a temporal-capable layer does not provide an override.temporal.layer_modes: optional object keyed bylayer_id; value enumannual_average|yearly|event.temporal.year_selection: optional global enumall|exclude_first|exclude_first_two|exclude_first_five|custom.temporal.exclude_yr_indxs: optional global array of zero-based integer indices.temporal.event: required when any effective temporal mode isevent.temporal.event.selector: enumdate|return_period.temporal.event.dates: required forselector=date; array ofYYYY-MM-DD.temporal.event.return_periods: required forselector=return_period; numeric array in years.tabular: optional object, valid only forformat=csv|parquet.tabular.concatenate_tables: optional boolean, defaultfalse.tabular.temporal_layout: optional enumwide|long, defaultwide.
Request example:
{
"format": "geoparquet",
"units": "project",
"crs": "wgs",
"layers": ["wepp.H.element.parquet", "wepp.chan.out.parquet"],
"output_scopes": ["baseline", "roads"],
"scenarios": ["thinned", "control"],
"swat_run_id": "latest",
"temporal": {
"mode": "yearly",
"layer_modes": {
"wepp.H.element.parquet": "yearly",
"wepp.chan.out.parquet": "annual_average"
},
"year_selection": "exclude_first_two",
"exclude_yr_indxs": [0, 1]
}
}
Validation:
scenariosandcontrast_idsare mutually exclusive.- Omni scenario and Omni contrast layer families cannot be mixed in one request.
- Omni scenario layers require
scenarios. - Omni contrast layers require
contrast_ids. - SWAT layers require a resolved
swat_run_id;latestis resolved to a concrete run ID before execution and persisted in manifest/cache key. - Unknown layer IDs return 400.
- Unsupported
crsvalue returns 400. - Unsupported temporal mode returns 400.
- Daily timeseries mode is not supported and returns 400.
swat_tables.includeandswat_tables.excludeare mutually exclusive.format=parquet|csvis valid for both spatial and non-spatial datasets and strips geometry from output rows instead of failing on spatial inputs.tabularis only valid whenformat=parquet|csv.tabular.concatenate_tablesmust be boolean when provided.tabular.temporal_layoutmust bewide|longwhen provided.tabular.temporal_layout=longrejects mixed effectiveeventandyearlylayer modes in one request.column_selection[layer_id].includeandcolumn_selection[layer_id].excludeare mutually exclusive.- Unknown layer ids in
column_selectionreturn 400 with structured validation errors. - Unknown column ids return 400 when the target layer has an explicit
columnscontract in catalog metadata; for discovery-driven layers without explicitcolumns, dynamic source-schema column ids are accepted. - If
column_selection[layer_id].includeis provided, exported fields for that layer are limited to the selected set plus required identity/join geometry fields. - If
column_selection[layer_id].excluderemoves all optional fields, export still retains required identity/join geometry fields. crs=utmrequires a resolvable run UTM CRS; unresolved UTM CRS returns 409.- AgFields WEPP metric layers require AgFields output/interchange assets; exporter performs on-demand preparation as defined in Section 6.3.
- Temporal mode support is evaluated per resolved layer from catalog
temporal.supported_modes. - Every selected temporal-capable layer must resolve an effective temporal mode from
temporal.layer_modes[layer_id]or fallbacktemporal.mode. - If
temporal.mode=yearly(or a layer-level effective mode isyearly) andtemporal.year_selectionis omitted, default toyear_selection=all. year_selectionandexclude_yr_indxsapply globally across all layers whose effective mode supports year filtering.- If some layers are incompatible with requested temporal settings and at least one layer remains exportable, incompatible layers are dropped with
layer_unavailablewarnings. - If no requested layers support the requested temporal settings, return 400.
- If
year_selectionorexclude_yr_indxsis provided for a layer whose catalog rule setsyear_selection_supported=false, ignore those selectors for that layer and emitselector_defaulted. - Missing required source dependencies for a resolved layer (missing required source locator, missing required source file, unsupported required source kind, unresolved required join key) fail the job with
materialization_error; silent downgrade to warnings is forbidden. - Optional missing datasets may emit warnings and still succeed when at least one export target resolves.
- Unsupported format dependency returns 409.
CRS behavior:
crs=wgsexports spatial layers in EPSG:4326.crs=utmexports spatial layers in the run-resolved UTM CRS (single resolved EPSG per job).- Non-spatial layers are unaffected by CRS selection.
- Geometryless formats (
parquet,csv) are unaffected by CRS selection because geometry is not exported.
Submission response:
- Always HTTP 202 with canonical async payload.
- Required key:
job_id. - Required
status_urlpoints to/rq-engine/api/jobstatus/{job_id}. - Required
download_urlpoints to/rq-engine/api/runs/{runid}/{config}/export/features/job/{job_id}/downloadand is only valid once the job isfinished. - Cache hits still return 202 with a new
job_id(fast-path job), never sync 200.
5.2 Polling And Result Contract
Polling is canonical RQ polling:
GET /rq-engine/api/jobstatus/{job_id}GET /rq-engine/api/jobinfo/{job_id}
Status semantics:
- Success terminal state is
finished. - Failure terminal states from job payload are
failed|stopped|canceled. - Job/status lookup misses are HTTP 404 error responses with
error.code="not_found". - Feature export does not define alternate terminal names like
completed.
Warnings and summaries:
jobstatuskeeps canonical fields (job_id,runid,status,started_at,ended_at).- Export warnings and manifest summary are carried in
jobinfo.result. jobinfo.resultminimum fields:artifact_iddownload_url(canonical job route URL)cache_hit(boolean)source_job_id(present on cache hit)manifest_relpathwarnings(array of warning objects)
5.3 Download
GET /api/runs/{runid}/{config}/export/features/job/{job_id}/download
GET /api/runs/{runid}/{config}/export/features/published/{profile}/download
Behavior:
- Job endpoint resolves
job_idto anartifact_idmapping. - Job endpoint returns file response when
jobstatus.status == "finished". - Job endpoint returns 409 if job is not yet terminal success.
- Job endpoint returns canonical 404 if job or artifact mapping does not exist.
- Published endpoint resolves
{profile}throughexport/features/published/index.json(source of truth). - Published endpoint profile tokens are canonical kebab-case profile IDs (for cutover:
prep-wepp,prep-wepp-geodatabase,prep-details); nolatestpath segment is used. - Published endpoint returns 404 when the profile has no published entry.
- Published endpoint returns 409
stale_publicationwhen the registry entry no longer maps to a valid cache/artifact binding for the published profile request. - Published endpoint sets
Content-Dispositionfilename as<runid>.<canonical-profile>.<format>.zip(for examplerun-1.prep-wepp.geopackage.zip).
6. Dependency Tracking And Options-Aware Caching
6.1 Execution Tracking Versus Cache Index
RedisPrep and cache index have separate responsibilities:
- RedisPrep + TaskEnum track transient workflow state and latest RQ job IDs.
- Persistent cache index tracks reusable artifacts by request+dependency fingerprint.
Required RedisPrep changes:
- Add
TaskEnum.run_features_export = "run_features_export"with labelExport Featuresand emoji📦. - Use
RedisPrep.timestamp(TaskEnum.run_features_export)for lifecycle milestones. - Persist latest export job ID under
RedisPrep.set_rq_job_id("features_export", job_id).
Persistent cache index:
- Store under run workspace at
export/features/cache/index.json. - Index key is
request_hash + dependency_fingerprint. - Index value includes
artifact_id, artifact paths, sourcejob_id, and manifest metadata.
6.2 Canonical Cache Key Rules
Cache key must use normalized payload and resolved dependencies:
- Normalize and sort arrays:
layers,output_scopes, table lists. - Resolve defaults before hashing (
crs,output_scopes,swat_run_id, temporal defaults). - Resolve
swat_run_id="latest"to concrete run id before hashing. - Include Unitizer settings fingerprint when
units=projectis used. - Include version markers: layer catalog version, unit conversion version, export code version.
- Include dependency fingerprint from the export dependency resolver (not RedisPrep/Preflight task status).
Dependency resolver contract:
- Resolve final dataset relpaths after all selectors are applied (
output_scopes, scenarios/contrast_ids, SWAT run/table filters, temporal mode). - Build dependency entries from actual resolved
geometry.locator,sources, anddependenciesinlayer_catalog.yaml, includingunitizer.nodbwhenunits=project. - Include
layer_catalog.yamlmetadata/version signature in dependency resolution. - Service fingerprints regular files from canonical relpath/provenance, size and verified SHA-256; mtime remains diagnostic. Explicit low-level metadata mode preserves size/mtime identity. See the file-content amendment below.
- Parent-run dependencies for canonical Omni child runs are valid cache dependencies when the resolved path stays within the inferred parent run root.
- Build the final dependency fingerprint from ordered identity projections serialized in canonical JSON; omit mtime only for valid SHA-256 entries.
6.3 AgFields Interchange Preparation (Parity+)
Trigger:
- Any requested layer in the AgFields WEPP metrics family.
- No backward-compatibility hooks are required for AgFields layer IDs or selectors; enforce the current parity+ contract directly.
Behavior:
- AgFields stage 4 publishes its six-file specialized interchange bundle synchronously before the RQ task stamps completion. A missing, stale, or version-incompatible bundle makes AgFields metric layers unavailable; export does not invoke the ordinary interchange migration path.
- The current metric layer joins sub-field geometry and PASS metrics on
sub_field_id. The native schemas carry bothsub_field_idandfield_idand do not expose the parent hillslopewepp_idortopaz_idas sub-field identity. The layer sources do not pre-joinfields.parquetor WAT: both PASS and WAT contain repeated temporal rows, so identity-only composition would be many-to-many, while the field mapping itself repeatsfield_id. A future WAT layer must declare its own daily temporal grain instead of being combined with event-grain PASS rows. The join explicitly allows repeated identity keys so temporal materialization can aggregate multiple events after the single geometry-to-PASS join. The existing draftag_fields.metrics.fieldscatalog ID is retained unchanged for compatibility but is not adapted to the specialized bundle in this package: attaching sub-field depths or event values directly to a whole-field polygon is scientifically misleading. A future field metric contract requires explicit area-weighted depth and summed-volume aggregation semantics. run_interchange_migration(..., "ag_fields")and ordinarytotalwatsed3.parquetare not valid preparation fallbacks because they apply ordinary watershed identity assumptions.- Required assets are catalog-driven from
ag_fields.metrics.subfieldsacrossgeometry.locator,sources, anddependencies; no hardcoded file list exists outside the catalog contract. - Submission performs a read-only readiness check before dependency and cache
planning. If any requested AgFields metrics layer lacks the current controller
completion marker or has a stale/incompatible bundle, reject the submission
with HTTP 409 and
ag_fields_interchange_not_current. The check does not run interchange, mutate project assets, or silently drop the requested layer from a mixed-layer export.
6.4 Cache Hit Behavior
Cache hit flow:
- Submit endpoint still enqueues a lightweight export-finalize RQ job and returns 202.
- Lightweight job writes a new job-scoped manifest that points to existing
artifact_id. jobinfo.result.cache_hit=trueandsource_job_id=<original producer job>.- Job download endpoint serves the cached artifact via
artifact_idmapping.
Artifact layout:
- Job metadata:
export/features/jobs/{job_id}/. - Reusable artifacts:
export/features/artifacts/{artifact_id}/. - Manifest exists in both locations.
- Job manifest includes
cache_hitandsource_job_id.
6.5 Published Profile Registry
Publication-level downloads use one run-scoped registry document:
- Path:
export/features/published/index.json. - This file is the source of truth for
GET /api/runs/{runid}/{config}/export/features/published/{profile}/download. - The registry is a lightweight JSON index; it must not be modeled as a dedicated NoDb controller class.
Registry contract:
- Top-level:
schema_version(integer),updated_at_utc(ISO 8601 UTC timestamp),profiles(object map keyed by canonical profile ID).
- Canonical profile IDs for legacy-cutover publication are
prep-wepp,prep-wepp-geodatabase, andprep-details. prep-wepp-gpkg-gdbis an execution-only virtual orchestration profile and is not persisted as a registry key; it co-publishesprep-weppandprep-wepp-geodatabase.- Each
profiles.{profile}entry includes:profile(string, matches key),job_id,artifact_id,artifact_relpath,manifest_relpath,format,request_hash,dependency_fingerprint,cache_key,published_at_utc.
- Registry writes are atomic and idempotent per profile key.
- Published download resolution must verify that the registry entry still maps to an existing artifact and a valid cache entry compatible with the canonical published profile format; registry fingerprint fields may be repaired from cache key components when recoverable.
6.6 WP-2 Milestone Status (Completed 2026-03-26)
ExecPlan completion:
docs/mini-work-packages/20260326_features_export_wp2_execplan.mdstatus isdone.
Implemented files:
wepppy/nodb/mods/features_export/dependency_tracker.pywepppy/nodb/mods/features_export/cache_key.pywepppy/nodb/mods/features_export/__init__.pytests/nodb/mods/test_features_export_dependency_tracker.pytests/nodb/mods/test_features_export_cache_key.py
Contract clarifications from implementation:
nodb_reflocator paths are resolved through an explicit resolver callback contract in WP-2 helpers; no implicit controller fallback behavior is used.path_templatelocators that include{table_name}require pre-resolved table names (for example SWAT table discovery output) before dependency fingerprinting.- Dependency fingerprints include a stable catalog metadata signature (
catalog_version,schema_version,updated_at_utc,owner,status) plus ordered canonical dependency entries. - Dependency entry snapshots include
relpath,exists,size,mtime_ns, and optionalcontent_hash_marker/content_hash_value(sha256mode). - Cache request hashing requires a concrete
swat_run_id(no unresolvedlatest) and requires Unitizer preferences fingerprint input whenunits=project. - WP-2 cache index helper persists deterministic JSON at
export/features/cache/index.jsonwith load/get/upsert semantics andschema_version=1.
Validation evidence:
wctl run-pytest tests/nodb/mods/test_features_export_dependency_tracker.py --maxfail=1-> pass (3 passed)wctl run-pytest tests/nodb/mods/test_features_export_cache_key.py --maxfail=1-> pass (4 passed)wctl run-pytest tests/nodb/mods/test_features_export_catalog_loader.py --maxfail=1-> pass (2 passed)
7. Units Strategy And Unitizer Requirements
Units modes:
si: SI export output.english: English export output.project: project Unitizer settings, including mixed per-variable preferences.
features_export must use Unitizer numeric conversion primitives.
Column naming contract:
- Unit-applicable output columns must include a normalized unit token suffix in the exported column name (for example
runoff_mm,hillslope_area_ha,runoff_volume_m3,sediment_yield_kg_m2). - Columns without an applicable unit mapping keep their canonical source name and are recorded as pass-through in manifest unit metadata.
- Manifest must include a per-column unit mapping table so UI/download consumers can recover source field, target field, and resolved unit metadata deterministically.
7.1 Unitizer Milestone Status (Completed 2026-03-25)
ExecPlan completion:
docs/mini-work-packages/20260325_unitizer_features_export_execplan.mdstatus isdone.
Implemented files:
wepppy/nodb/unitizer.pywepppy/nodb/unitizer.pyitests/nodb/test_unitizer_numeric_apis.pydocs/mini-work-packages/20260325_unitizer_features_export_execplan.md
New public Unitizer API contract:
- Numeric conversion APIs:
convert_scalar,convert_sequence,convert_table. - Target resolution API:
resolve_target_unitwithsi|english|projectsemantics. - Stable cache fingerprint API:
preferences_fingerprint. - Public metadata/result types:
UnitTargetResolution,UnitConversionMetadata,UnitizedScalar,UnitizedSequence,UnitizedTable. - Public helper:
get_unit_class. - Explicit pass-through/no-mapping signaling via
pass_through_reason. - Ambiguity handling for shared labels (for example
ppm) and identity-path type preservation (no int-to-float coercion when no conversion applies). - Existing
context_processor_package()behavior remains compatible (unitizer,unitizer_units,unitizer_with_units).
Validation evidence from handoff:
wctl run-pytest tests/nodb/test_unitizer_preferences.py --maxfail=1-> pass (4 passed)wctl run-pytest tests/weppcloud/routes/test_unitizer_bp.py --maxfail=1-> pass (2 passed)wctl run-pytest tests/nodb/test_unitizer_numeric_apis.py --maxfail=1-> pass (29 passed)wctl run-stubtest wepppy.nodb.unitizer-> passwctl check-test-stubs-> passwctl run-pytest tests --maxfail=1-> pass (2582 passed, 34 skipped)
Reviewer status:
- High/medium findings reported during review were resolved (ambiguity handling and identity conversion type preservation).
- QA reviewer reported no remaining high/medium findings.
Residual risk:
- Low-risk untested defensive branch remains (
target_unit_not_supported), requiring registry mutation to exercise. - No blocking risks identified.
8. Temporal Semantics
Temporal schema policy:
- Preserve native source schemas and temporal grain.
- For
eventandyearlyexports, materialize temporal measures in wide form at carrier geometry grain (one spatial feature row per canonical key).
Supported modes:
annual_averageyearlyevent
Global temporal control model:
- Temporal mode is resolved per selected temporal-capable dataset from
temporal.layer_modes[layer_id]with fallback totemporal.mode. - Global year selection controls (
year_selection,exclude_yr_indxs) apply across all datasets whose effective temporal mode supports year filtering. - UI control order must place global year selection immediately after CRS selection.
- Per-dataset column selection is independent from temporal controls and applies after temporal filtering.
annual_average rules:
- Uses return-period year-selection behavior.
exclude_yr_indxsuses zero-based year index semantics consistent with return-period processing.year_selection=customrequires explicitexclude_yr_indxs.
yearly rules:
- If
year_selectionis omitted, default toall. - Export includes every available year after global year filters are applied.
yearlywide materialization pivots selected measures to year-suffixed columns (for examplerunoff_yr2015_mm).- When multiple rows exist for one
{key, year}slice, numeric measures are explicitly reduced by summation before pivoting; conflicting non-numeric slices fail withmaterialization_error. - If year filtering excludes all years for a layer, that layer is dropped with
layer_unavailable(or 400 if no layers remain).
event rules:
selector=date: explicit date set.selector=return_period: explicit requested recurrence intervals (years), never raw rank values.- Mixed date and return-period selectors in one request are invalid.
- Event selector filtering is applied to discovered source frames before key-first uniqueness checks.
- Required sources missing selector-compatible columns must fail with
materialization_error. - Required sources with selector-compatible columns but zero matched rows remain materialized as empty event cores; export succeeds with canonical geometry rows and null event metrics.
- Optional sources that cannot satisfy the active event selector are skipped.
- Return-period filtering uses nearest available Weibull
TwhereT >= requested_interval(one available interval may satisfy at most one requested interval). - Rank-only lookup sources derive
Tusing the canonical WEPP return-period defaults (method=cta, Gringorten correction enabled) before applying interval matching. - Event materialization pivots selected measures to selector-token-suffixed columns (for example
q_2015_01_16_mmorrunoff_rp2_mm) so output geometry remains normalized to canonical carrier feature counts. - If event slices contain duplicate OFE rows, the terminal OFE (
max(ofe_id)) is selected as the deterministic per-slice representative before pivoting. - Remaining conflicting duplicates for the same
{key, event_token, measure}slice are contract failures (materialization_error).
Mixed-layer temporal compatibility:
- Temporal compatibility is layer-specific and driven by each layer's catalog
temporal.supported_modesandtemporal.mode_rules. - Layers with
temporal.supported_modes=[]are explicitly atemporal and remain exportable regardless of request temporal mode. - Layers incompatible with request temporal selectors are excluded with
layer_unavailablewhen at least one other layer remains exportable. - If every requested layer is excluded by temporal compatibility checks, return 400.
year_selectionandexclude_yr_indxsare only applied where catalog rules allowyear_selection_supported=true; otherwise emitselector_defaulted.
9. Selector Rules For Omni And SWAT
Omni:
- Scenario layers require
scenarios. - Contrast layers require
contrast_ids. - Scenario and contrast families cannot be requested together in one job.
- Omni contexts inherit the selected base WEPP datasets; users do not pick separate Omni dataset lists.
- Omni selectors are multi-select and support bulk controls (
Select All,Unselect All) for discovered options (required for contrasts).
SWAT:
- Default is all discovered interchange tables for resolved
swat_run_id. swat_tables.includeexports only listed tables.swat_tables.excludeexports all discovered minus listed tables.- Include/exclude values are deduplicated and lexicographically sorted for cache canonicalization.
- SWAT table resolution is profile-driven from catalog
table_profiles(for example,subbasin,channel,hru,non_spatial). - Each resolved table maps to profile-defined geometry strategy and join contract before export.
- Non-spatial SWAT tables are exportable for
geoparquet,parquet,csv,geopackage, andgeodatabase; they are skipped withtable_unavailablewarnings forgeojsonandkmz.
10. NoDb Mod And Runs-Page UI Contract
Module placement:
- Implement at
wepppy/nodb/mods/features_export. - Follow NoDb facade/collaborator pattern.
Runs page integration:
- Add
Exportto the Mods list. - Implement a NoDb controller UI using established async pattern.
- Add top-of-control profile actions:
Load Export Profilequick actions are populated from built-in profile files discovered viaload_builtin_profiles()plus virtual orchestration presets (current built-ins:Prep details,Post Wepp,Temporal yearly; virtual:Post Wepp (GPKG + GDB)/prep_wepp_gpkg_gdb).Specify Export from Profiletext area +Load profileaction.Clear selectionremains available as a separate action.
Post Weppis the default quick profile and replaces the legacyLoad Defaultsbutton behavior.- Virtual profile discoverability contract:
prep_wepp_gpkg_gdbis emitted in runs-page bootstrapprofiles/profile_buttonseven though it has no dedicated.ymlfile.- Its base request is resolved via
resolve_published_profile_request("prep-wepp-gpkg-gdb"). - Runtime enrichment applies before exposing it to the UI:
- add
roadstooutput_scopeswhen roads scope is available for the active run/config; - add
omni.scenarios.hillslopesand discovered scenario IDs (scenarios) when Omni scenarios are available.
- add
- Profile text loading accepts pasted YAML/JSON request-profile content and applies the profile without auto-submit.
- Run settings visual order is fixed:
format->units->crs-> globalyear_selection. - Catalog UI is hierarchy-first and must not include a layer search box, filter chips, or "select visible" behavior.
- Family labels are user-facing domain labels and must use one consolidated
WEPPfamily with familiar output names (not splitWEPP Summary,WEPP Temporal,WEPP Interchangeheadings). - Layer rows must present clear hierarchy/indentation under family headers rather than a flat left-aligned list.
- Each dataset row must include an expandable/collapsible "Columns" section showing:
- Column checkbox (selected/unselected)
- Column label /
column_id - Source-backed description text when available
- Resolved unit display (or explicit non-unitized marker)
- Required-field indicator for non-removable identity/join columns
- The collapsed row remains scannable; detailed column picking is opt-in through expansion.
- Column metadata source order is: parquet field metadata (
label,description,units) -> interchangeREADME.mddocs for the resolved source file -> deterministic fallback label/unit inference. - Required identity/join locks are canonicalized by column token so alias-equivalent keys (for example
topaz_idvsTopazID) do not render as duplicate mandatory selectors. - Every temporal-capable dataset row includes a temporal mode control (dataset-scoped mode); global year selection remains single and shared.
- Omni Scenarios and Omni Contrasts families render at the bottom of the catalog.
- Output scope controls are discovery-aware: disable
roadswith an explanatory hint when roads outputs are unavailable for the active run/config. - Family discovery is dynamic: hide groups with no available datasets (for example AgFields when missing inputs).
- Availability/scope readiness updates should stream through websocket status updates so users do not need a manual dataset-detection action.
- Use the dedicated subagent role pack at
wepppy/nodb/mods/features_export/SUBAGENT_ROLES.mdfor UI design/development specification and implementation planning. - Use
wepppy/nodb/mods/features_export/ui_control_layout.mdas the canonical detailed control layout and ASCII wireframe reference. - Controller posts JSON payload, stores returned
job_id, and polls canonical/rq-engine/api/jobstatus/{job_id}viaset_rq_job_id. - Completion details and warnings are read from
/rq-engine/api/jobinfo/{job_id}. - Download is enabled when job state is
finished. - Controller must attach
attach_status_streamwith stacktrace hooks and keep poll fallback enabled. - Controller must hydrate prior
job_idon bootstrap using existing controller-contract guidance. - Template must include required status panel, stacktrace panel, and job-hint DOM hooks with
aria-live="polite"status behavior. - Built-in profile source-of-truth files live in:
wepppy/nodb/mods/features_export/profiles/post-wepp.ymlwepppy/nodb/mods/features_export/profiles/prep-details.ymlwepppy/nodb/mods/features_export/profiles/temporal-yearly.yml
- Virtual quick profiles (for example
prep_wepp_gpkg_gdb) are discoverable through bootstrap payload composition and intentionally do not require a dedicated file underprofiles/. - Published download profile IDs are canonical kebab-case tokens (
prep-wepp,prep-wepp-geodatabase,prep-details) and may map to built-in profile aliases during cutover (post_wepp->prep-wepp,prep_details->prep-details). prep-details.ymlis the canonical replacement profile for legacyprep_detailsexport behavior and defaults toformat=csv.temporal-yearly.ymlis the canonical built-in preset that exercises yearly temporal measures (wepp.interchange.loss_all_years_hill).
11. Manifest And Warning Contract
Every artifact includes:
manifest.json(canonical machine-readable metadata/provenance contract).- generated
README.md(human-readable metadata summary derived from resolved export metadata).
11.1 Geospatial Metadata Standards Baseline For Artifact README
The generated README.md must align with established geospatial metadata guidance and format standards:
- FGDC CSDGM v2 (
FGDC-STD-001-1998) and FGDC CSDGM Essential Metadata Elements for minimum discovery, contact, extent, quality, and lineage coverage. - FGDC-endorsed ISO 191** metadata suite baseline with ISO 19115-1 Fundamentals as the core model.
- USGS metadata best practices for practical quality guardrails (descriptive title, abstract/purpose, update dates, DOI-as-URL when present, and packaging metadata with data).
- GeoParquet v1.1.0 metadata requirements for geometry encoding and CRS semantics in parquet payloads.
- Reference: https://geoparquet.org/releases/v1.1.0/
- OGC GeoPackage metadata extension model (
gpkg_metadata,gpkg_metadata_reference) for standards-compatible metadata carriage in GPKG ecosystems. - GeoJSON RFC 7946 CRS and bbox semantics for GeoJSON payload interpretation.
- Reference: https://www.rfc-editor.org/rfc/rfc7946
11.2 Available Metadata Inputs For README Generation (Current Implementation)
Metadata already available today (no new science/data-source contracts required):
- Manifest/request context:
request.resolved(format,units,crs,output_scopes, temporal selectors, scenario/contrast selectors, SWAT selectors, column selection, tabular options)generated_at_utc,cache_hit,source_job_id,artifact_id
- Artifact and packaging context:
artifact.format,artifact.artifact_relpath,artifact.packaged_member_relpaths
- Layer-level context:
layer_id,output_layer_id,family,scope_class,scope,context,selector_id,carrier_layer,temporal_moderow_count,feature_count,artifact_relpath- output column metadata (
source_layer_ids,selected_columns,unit_mapping,description_mapping, materialization strategy metadata)
- CRS/projection context:
crs.requested_crs,crs.resolved_crs,crs.resolved_epsg(when available)
- Dependency and lineage context:
dependency_snapshot.catalog_signature,dependency_snapshot.fingerprint- per dependency entry:
relpath,exists,size,mtime_ns,content_hash_*,dependency_role,dependency_id
- QA/status context:
warningswith canonical warning codes and optionallayer_id/scope
Known metadata gaps to track separately (do not block initial README rollout):
- Persistent identifiers (DOI/PID) for exported artifacts.
- Explicit distribution license/use constraints and access constraints per export artifact.
- Canonical contact/organization fields for artifact-level metadata ownership.
- Spatial extent (
bbox) and temporal extent summaries precomputed across all exported layers. - Formal data-quality measure blocks (beyond warning summaries and source dependency fingerprinting).
11.3 Dynamic Artifact README Contract
README generation behavior:
README.mdis generated dynamically for each cache-miss artifact publication and packaged into the artifact zip root.- Cache-hit jobs reuse the immutable artifact
README.mdand bundled manifest. Job-scoped manifests may record the new job timestamp,cache_hit/source_job_id, current dependency observation (including diagnostic mtime), and selection verification context. They must preserve the artifact identity and accepted content fingerprint; new observations do not rewrite producer provenance. README.mdis deterministic for the same artifact payload/manifest inputs (stable ordering and section structure).
README minimum sections:
- Export summary:
- generated timestamp, run/config context, format, units mode, CRS mode.
- Standards and interpretation notes:
- concise format-specific CRS/metadata interpretation notes (for example GeoJSON RFC 7946 WGS84 semantics, GeoParquet CRS notes).
- Resolved request profile:
- normalized selectors (
layers,output_scopes, temporal selectors, scenario/contrast selectors, SWAT selectors, tabular layout controls).
- normalized selectors (
- Layer inventory table:
- output layer id, source layer id(s), context/scope, row count, feature count, artifact member path.
- Column and unit summary:
- selected columns, resolved unit mapping, and column descriptions per output layer.
- Dependency lineage summary:
- dependency fingerprint, catalog signature, and grouped dependency entries by role.
- Warning summary:
- warning code/message table with layer/scope attachments when present.
- Machine-readable contract pointer:
- explicit pointer that
manifest.jsonis the canonical machine-readable provenance payload.
- explicit pointer that
README authoring rules:
- Do not include absolute host filesystem paths.
- Do not include secrets/tokens/auth headers.
- Avoid profile replay payload embedding (
profile.ymlis not bundled). - Prefer concise tables and stable ordering to keep diffs/cache artifacts deterministic.
Manifest minimum fields:
- Resolved request payload and selector defaults.
- CRS metadata (
requested_crs,resolved_crs,resolved_epsg). - Resolved dependency entries with path, existence, timestamp, and fingerprint components.
- Per-layer scope metadata (
baseline|roads|shared). - Layer context metadata (
base|scenario|contrast), selected selector id when applicable, and consolidated geometry carrier (sbs_map-subcatchments|chan_map-channels). - SWAT table profile resolution and per-table spatiality classification.
- Temporal compatibility decisions (selectors applied, selectors defaulted, and layer/table exclusions).
- Conversion summary and unit pass-through fields.
- Unitized column-name mapping (
source_column,export_column,resolved_unit,pass_through_reason). - Column-selection decisions by layer (
include,exclude, and required columns auto-retained). - Row and feature counts per layer.
- Generation timestamps and tool/catalog versions.
cache_hit,source_job_id,artifact_id.- Dependency-preparation records (including AgFields interchange auto-prep attempts and outcomes).
warningsarray.- Optional publication metadata when a job is promoted to published profile status:
published_profile,published_at_utc.
Warning object shape:
code: machine-readable warning code.message: human-readable description.layer_id: optional associated layer.scope: optional associated scope.
Reserved warning codes:
scope_missing_layerscope_not_applicablelayer_unavailabletable_unavailablemeasure_unavailableunit_pass_throughselector_defaultedroads_scope_unavailablelegacy_flags_ignored
12. Migration And Cutover
Cutover is immediate with explicit legacy cleanup:
- Remove direct
gpkg_exportroute/task usage from rq-engine export routes. - Remove direct
prep_detailsroute/task usage from rq-engine export routes. - Move export ownership from legacy modules (
wepppy/export/gpkg_export.py,wepppy/export/prep_details.py) to NoDbfeatures_export. - Rewire run-completion hooks to
features_exportprofile execution/publication. - Keep
/export/geopackage,/export/geodatabase, and/export/prep_detailsas compatibility facades that executefeatures_exportprofiles. - Standardize job downloads on
/export/features/job/{job_id}/download(replace/export/features/{job_id}/download). - Add profile-aware published downloads on
/export/features/published/{profile}/downloadbacked byexport/features/published/index.json. - Keep run-completion toggles functional while routing generation through
features_exportinternals. - Keep artifact bundles profile-file-free (
profile.yml, built-in profile files are excluded) while including generated artifactREADME.mdplusmanifest.json. - AgFields parity+ support ships without legacy compatibility shims (single-project assumption).
Back-compat behavior for existing saved configs:
- Persisted run-completion export flags remain active but now drive
features_exportprofile execution:prep_details_on_run_completion-> publishedprep-details.arc_export_on_run_completion-> published orchestration profileprep-wepp-gpkg-gdb(co-publishesprep-wepp+prep-wepp-geodatabase).
- Legacy module imports/writers are removed; flags do not call
wepppy/export/gpkg_export.pyorwepppy/export/prep_details.py. - Legacy module deletion (
gpkg_export.py,prep_details.py) must occur only after explicit human approval based on parity validation evidence (see work-package gate requirements).
13. Acceptance Criteria
- All seven formats export successfully on representative runs.
- Dataset merge/materialization path is DuckDB-first for production export payload assembly (no pandas merge loops on the hot path).
- Materialization is key-first/geometry-last: exactly one DuckDB carrier core table per
{context, selector_id, scope, carrier}plus one final geometry attach. - Export row counts are bounded by carrier key cardinality; multiplicative row growth from repeated many-to-many joins is a contract failure.
- Single-layer formats produce zipped files with one file per resolved layer.
- Multi-layer formats produce one container artifact per request.
- Base WEPP context exports at most two consolidated spatial layers per requested scope (
sbs_map-subcatchmentsand/orchan_map-channels). - Each selected Omni scenario/contrast exports its own consolidated
subcatchmentsand/orchannelslayers per requested scope. - Consolidated layer names follow descriptive run/context naming (for example
{runid}-roads-sbs_map-subcatchments). - Partial-scope exports emit warnings and still succeed when at least one scoped layer resolves.
- Non-scope missing requested layers/tables emit
layer_unavailableortable_unavailablewarnings and still succeed when at least one export target resolves. - Geometryless formats (
parquet,csv) export tabular outputs with geometry removed while preserving required identity/join columns. - Geometryless formats expose
tabular.concatenate_tablesandtabular.temporal_layoutcontrols with deterministic writer behavior. - Parity+ AgFields support is present: boundaries/sub-fields plus AgFields WEPP metric layers sourced from
wepp/ag_fields/output/interchange/*. - Requesting AgFields WEPP metric layers triggers on-demand AgFields interchange preparation when needed and proceeds without manual pre-run migration.
- Submit endpoint rejects non-JSON payloads with 415 and validates selector rules with 400/404/409 per contract.
- CRS selection works with
wgs|utm, defaults towgs, and exports include CRS metadata in manifest. - UI places global year selection directly after CRS selection.
- Temporal-capable datasets expose per-dataset temporal mode controls.
- Dataset rows expose expandable column-selection panels with visible units.
- Users can include/exclude columns per selected dataset; required identity/join columns remain enforced.
yearlymode exports every available year by default (year_selection=all) unless globally filtered.- Global year selection applies consistently across all applicable selected datasets.
eventandyearlyexports default towidetemporal layout with temporal selector tokens appended to measure columns.tabular.temporal_layout=longemits one temporal selector column (date,return_period, oryear) and rejects mixedevent+yearlyrequests.- Polling and terminal status behavior align with canonical RQ contract (
finishedsuccess). - Warnings are present in both
jobinfo.result.warningsandmanifest.json. - Cache hits return 202 with a new lightweight job ID,
cache_hit=true, and reusable artifact delivery. - Cache key includes Unitizer settings fingerprint for
units=projectand invalidates on Unitizer preference changes. - Dependency fingerprint is computed from resolved geometry/sources/dependencies, not task timestamps/preflight status.
output_scopesis case-normalized, deduplicated, and invalid values fail with 400.layer_catalog.yamlis present, machine-readable, and drives layer discovery/validation.- UI catalog shows one consolidated
WEPPfamily with familiar output names (for exampleH.element.parquet) and does not surface internal labels such aswepp.temporal.events. - Omni Scenarios and Omni Contrasts render at the bottom of the catalog.
- Omni selectors support multi-select and bulk selection controls (
Select All,Unselect All). - Layer search/filter strip is removed from the control.
- Discovery hides unavailable families and disables unavailable
roadsscope without requiring manual detection refresh. layer_catalog.yamlschema validation enforces locator vocabulary, temporal mode rules, and join-key precedence.- Optional measure availability (for example
tsmf, phosphorus,QRain,QSnow) follows catalog rules and emitsmeasure_unavailablewarnings when absent. - Baseline default export for run
clogging-starch/disturbed9002-wbt-mofeis a regression anchor with exactly two spatial layers and carrier-aligned feature counts (66subcatchments,27channels). - RedisPrep/TaskEnum timestamps and
rq:features_exporttracking are wired. units=si|english|projectall work with manifest conversion metadata.- Unit-applicable exported columns include unit tokens in column names and are documented in manifest column mapping metadata.
- Applied per-layer column-selection decisions are recorded in manifest metadata.
- Unitizer numeric API foundation (
convert_scalar|convert_sequence|convert_table,resolve_target_unit,preferences_fingerprint) is complete and available forfeatures_exportintegration. annual_average,yearly, andeventmodes work; daily is rejected.- Mixed temporal requests have deterministic partial-export behavior and warning semantics.
- SWAT default all-table export and include/exclude overrides work with profile-based geometry/non-spatial handling.
- Export mod appears on Runs page and follows existing NoDb async controller interaction.
- Run-page control exposes profile quick buttons (current set: built-ins
Prep details,Post Wepp,Temporal yearly, plus virtualPost Wepp (GPKG + GDB)) and a profile-text load path for pasted profile content. - All artifact downloads are zip bundles and include export payload members plus
manifest.jsonand generatedREADME.md. - Generated
README.mdincludes standards-aligned metadata sections (summary, layer inventory, CRS/units, dependency lineage, warnings, and manifest pointer). - Download routes are standardized to
/export/features/job/{job_id}/downloadand/export/features/published/{profile}/download. export/features/published/index.jsonis authoritative for published profile download resolution (prep-wepp,prep-wepp-geodatabase,prep-details).- Regression tests cover payload shape, selector validation, scope behavior, cache hit flow, and legacy cutover.
14. Implementation Skeleton And Work-Package Breakdown
This feature should be implemented as multiple small, ordered work-packages with stable boundaries and explicit handoff contracts.
Code organization target:
wepppy/nodb/mods/features_export/
__init__.py
specification.md
ui_control_layout.md
layer_catalog.yaml
contracts.py # request/result/warning dataclasses and enums
facade.py # NoDb facade entrypoint for control/bootstrap data
catalog_loader.py # catalog read + schema validation + index build
planner.py # selectors -> resolved layer export plan
dependency_tracker.py # RedisPrep/TaskEnum preflight/dependency checks
cache_key.py # options-aware canonical cache key/fingerprint
unit_conversion.py # Unitizer numeric conversion integration adapter
column_selection.py # selected-column resolution and unit inference helpers
cache_rehydration.py # cache-entry artifact/layer-output rehydration helpers
discovery.py # source discovery and strict required-source handling
join_planner.py # join-key normalization and cardinality contracts
duckdb_materializer.py # key-first DuckDB attribute materialization
geometry_carriers.py # canonical geometry carrier build + final attach
manifest_builder.py # manifest-facing per-layer column metadata assembly
legacy_source_materializer.py # legacy geometry-first source merge collaborator
carrier_layer_materializer.py # carrier-source materialization collaborator
exporters/
__init__.py
base.py # common writer contract
geojson.py
geoparquet.py
parquet.py
csv.py
kmz.py
geopackage.py
geodatabase.py
packaging.py # zip bundling for single-layer formats
manifest.py # manifest/warning assembly
service.py # orchestration: plan -> export -> manifest -> artifact
Related integration files (outside module) should be kept thin and delegate to features_export.service:
- rq-engine export route/controller.
- rq task enqueue + worker function.
- Runs-page controller bootstrap + template wiring.
14.1 Work-Package Sequence
WP-0: Unitizer numeric API foundation (completed)
- Status: done via
docs/mini-work-packages/20260325_unitizer_features_export_execplan.md. - Deliverable: public numeric conversion APIs + cache fingerprint support now available for
features_export.
WP-1: Core contracts and planner skeleton (completed 2026-03-26)
- Status: done via
docs/mini-work-packages/20260326_features_export_wp1_execplan.md. - Implemented files:
contracts.py,catalog_loader.py,planner.py,__init__.py, and focused tests undertests/nodb/mods/test_features_export_*. - Contract clarification: planner validation failures are emitted through a canonical 400-style
validation_errorcontract (error + errors[]) with structuredValidationIssueentries. - Contract clarification: deterministic plan ordering is canonicalized by sorted unique
layers, canonicaloutput_scopesorder (baseline,roads), and stable resolved layer ids ({scope_or_shared}__{layer_id}). - Contract clarification: temporal incompatibility is handled per layer; incompatible layers emit
layer_unavailable, and if all layers are excluded the planner raisesno_exportable_layers. - Deliverable: deterministic
ResolvedExportPlanobject and unit tests for selector rules.
WP-2: Dependency tracking and options-aware caching (completed 2026-03-26)
- Status: done via
docs/mini-work-packages/20260326_features_export_wp2_execplan.md. - Implemented files:
dependency_tracker.py,cache_key.py, package export wiring, and focused tests undertests/nodb/mods/test_features_export_dependency_tracker.pyandtest_features_export_cache_key.py. - Contract clarification: dependency locators enforce strict
kind/valuestructure at resolution time;nodb_refrequires explicit resolver input andpath_templateentries requiring{table_name}require pre-resolved table names. - Contract clarification: request hash requires concrete
swat_run_idand includes Unitizer preferences fingerprint forunits=project, plus catalog/conversion/export version markers. - Deliverable: deterministic dependency snapshot/fingerprint and deterministic cache-key/index foundation (
export/features/cache/index.json).
WP-3: Format writers, packaging, and manifest generation (completed 2026-03-26)
- Status: done via
docs/mini-work-packages/20260326_features_export_wp3_execplan.md. - Implemented files:
wepppy/nodb/mods/features_export/exporters/__init__.pywepppy/nodb/mods/features_export/exporters/base.pywepppy/nodb/mods/features_export/exporters/geojson.pywepppy/nodb/mods/features_export/exporters/geoparquet.pywepppy/nodb/mods/features_export/exporters/kmz.pywepppy/nodb/mods/features_export/exporters/geopackage.pywepppy/nodb/mods/features_export/exporters/geodatabase.pywepppy/nodb/mods/features_export/exporters/packaging.pywepppy/nodb/mods/features_export/manifest.pytests/nodb/mods/test_features_export_exporters.pytests/nodb/mods/test_features_export_manifest.py
- Contract clarification: writer inputs are explicitly pre-resolved and payload-driven (
ResolvedExportPlan+ per-layerPreparedLayerPayloadmapping keyed byoutput_layer_id); WP-3 does not resolve or extract source datasets. - Contract clarification: single-layer formats (
geojson|geoparquet|kmz) write deterministic one-file-per-layer outputs and return one deterministic zip bundle; multi-layer formats return one container artifact per request. - Contract clarification: geodatabase writer directly invokes bounded
ogr2ogr -f OpenFileGDBconversion in the worker runtime and fails explicitly when the built-in driver lacks vector-create capability. - Contract clarification: geopackage artifacts must be valid SQLite/GPKG containers (not synthesized JSON payload bytes); geodatabase staging gpkg input must use the same container contract.
- Contract clarification: geopackage writer output must be GDAL/OGR-readable for downstream conversion boundaries; implementation uses the GDAL
GPKGdriver and supports both spatial feature payloads and aspatial fallback payloads for interoperability. - Contract clarification: feature-layer field synthesis must retain null-only properties as nullable columns while preserving numeric field typing for non-null numeric properties.
- Contract clarification: manifest assembly is pure (
build_export_manifest) and serialization/write (serialize_export_manifest,write_export_manifest) is a separate step. - Deliverable: deterministic artifact metadata and manifest generation from pre-resolved plan inputs.
WP-4: Service orchestration and RQ wiring
- Add
service.pyorchestration and thin rq-engine route/task adapters. - Implement canonical async submit/jobinfo/download behavior.
- Deliverable: end-to-end export job execution for API callers with warning contract.
WP-5: Runs-page UI control integration
- Status: completed 2026-03-26 via
docs/mini-work-packages/20260326_features_export_wp5_execplan.md. - Implemented files:
wepppy/weppcloud/routes/run_0/run_0_bp.pywepppy/weppcloud/routes/run_0/templates/runs0_pure.htmwepppy/weppcloud/routes/run_0/templates/run_page_bootstrap.js.j2wepppy/weppcloud/templates/header/_run_header_fixed.htmwepppy/weppcloud/controllers_js/project.jswepppy/weppcloud/templates/controls/features_export_pure.htmwepppy/weppcloud/controllers_js/features_export.jswepppy/weppcloud/controllers_js/__tests__/features_export.test.jswepppy/weppcloud/static-src/tests/smoke/controller-cases.jstests/weppcloud/routes/test_pure_controls_render.pytests/weppcloud/routes/test_project_bp.pytests/weppcloud/routes/test_run_0_openet_admin_gate.py
- Contract clarification: Runs-page dynamic mod behavior requires both server-side mod metadata (
MOD_UI_DEFINITIONS+view/mod/<mod_name>) and client-side bootstrap registration (project.jsMOD_BOOTSTRAP_MAP) for runtime mod insertion parity with initial page render. - Contract clarification: the features-export submit route is
rq:exportscoped and explicitly documents/requires415for non-JSON payloads; frozen checklist/rules artifacts were aligned accordingly. - Contract clarification: smoke-case execution requires a pre-submit layer selection for
features_exportbecause the form submit action is validation-gated until minimum payload requirements are met. - Contract clarification: service orchestration must always pass a
nodb_ref_resolverinto dependency snapshot construction so catalognodb_reflocators (for examplenodb:watershed.subwta_shp) resolve deterministically during submit-time cache/dependency planning. - Contract clarification: features-export status UI must use canonical
control_shellstatus-panel plumbing (status_panel_options) sowc-status-paneltheming and shared status-log behavior remain consistent with other controllers. - Contract clarification: controller submit/bootstrap paths must resolve job IDs from canonical variants (
job_id, wrappedContent.job_id, and keyedjob_idsmaps includingrun_features_export/run_features_export_rq) before treating submit as failed. - Contract clarification: completed features-export result payloads should provide dedicated rq-engine download links (
/api/runs/{runid}/{config}/export/features/job/{job_id}/download) instead of browse-service relpath URLs. - Contract clarification: cache-hit reuse must validate geopackage artifact signature; legacy non-SQLite
.gpkgcache entries are treated as invalid and regenerated through cache-miss execution. - Contract clarification: rq-engine enqueue selection for cache-hit worker must use validated cache eligibility (artifact format/path integrity), not cache-index presence alone, so invalid legacy entries route to standard execution.
- Contract clarification: submit-time service payload preparation must materialize catalog/dependency-resolved source data into per-layer feature collections (geometry + joined attributes) for default exports; payload-metadata-only
.gpkgrows are non-compliant for the Runs-page experience. - Deliverable: fully wired Runs-page control with Jest coverage and updated smoke/route-template invariants.
WP-6: Cutover and legacy retirement
- Status: complete via
docs/work-packages/20260329_features_export_legacy_exports_cutover/(including GO-approved Phase 8 legacy module deletion). - Active planning package:
docs/work-packages/20260329_features_export_legacy_exports_cutover/. - Replace legacy geopackage/prep-details routes and completion hooks with
features_exportprofile-backed execution. - Add publication-aware download path and registry (
/export/features/published/{profile}/download,export/features/published/index.json) with canonical published profilesprep-wepp,prep-wepp-geodatabase, andprep-details. - Remove compatibility
/export/features/{job_id}/downloadroute in favor of/export/features/job/{job_id}/download. - Contract clarification (2026-03-29): rq-engine legacy endpoints
/export/geopackage,/export/geodatabase, and/export/prep_detailsexecute throughresolve_published_profile_request+execute_features_export;geopackageandprep_detailspublish toexport/features/published/index.json. - Contract clarification (2026-03-29): geodatabase cutover path prefers published
prep-wepp-geodatabaseartifact resolution and only triggers on-demand execution when absent/stale. - Contract clarification (2026-03-29): post-WEPP completion hooks
_post_gpkg_export_rqand_post_prep_details_rqexecute profile-backed features-export flows;_post_gpkg_export_rqruns orchestration profileprep-wepp-gpkg-gdbthat executesprep-wepponce, co-creates.gdb, and publishes bothprep-weppandprep-wepp-geodatabase. legacy_flags_ignoredwarning code remains reserved for explicit legacy-flag migration messaging but is not required for normal profile-routed run-completion execution.- Require explicit human approval gate after parity and e2e validation evidence before deleting legacy modules (
wepppy/export/gpkg_export.py,wepppy/export/prep_details.py). - Deliverable: legacy replacement shipped with parity evidence, human approval record, and legacy modules removed.
WP-7: Reconciliation pass for WEPP naming, temporal controls, and consolidated layer outputs (planned 2026-03-27)
- Status: complete via
docs/mini-work-packages/20260327_features_export_reconciliation_execplan.md; architecture correction landed in WP-8. - Scope: reconcile taxonomy/UI/selector behavior with operator expectations, replace merge hot path with DuckDB-oriented consolidation for WEPP/Omni contexts, and land deterministic naming/temporal contracts.
- Contract clarification: WEPP outputs are presented as one family with familiar output names; internal labels (for example
wepp.temporal.events) are hidden. - Contract clarification: yearly mode must export all years by default and year-selection controls are global while temporal mode selection is dataset-scoped.
- Contract clarification: Omni scenario/contrast selection is multi-select, inherits base WEPP output selection, and supports
Select All/Unselect All. - Contract clarification: base and Omni WEPP outputs are consolidated to up to two geometry-carrier layers per scope with descriptive run/context layer naming.
- Contract clarification: control layout is hierarchy-first, removes search/filter strip, hides unavailable families, and receives websocket-driven discovery updates.
- Contract clarification: geometryless format tokens
parquetandcsvare added as first-class single-layer tabular exports that drop geometry while keeping required identity/join columns. - Contract clarification: profile-based defaults are first-class (
post-wepp.yml,prep-details.yml,temporal-yearly.yml) and replace the legacy defaults/prep-details split behavior. - Deliverable: reconciled backend/UI contract with regression and performance validation coverage.
WP-8: Key-first carrier materialization rewrite and module maintainability refactor (planned 2026-03-27)
- Status: complete via
docs/mini-work-packages/20260327_features_export_key_first_materialization_execplan.md(validated 2026-03-28 with66/27baseline carrier counts onclogging-starch/disturbed9002-wbt-mofe). - Scope: replace geometry-first dataset materialization with key-first DuckDB carrier-core assembly, enforce one-row-per-key join contracts, canonicalize carrier geometry, and split service orchestration into maintainable collaborators (
discovery,join_planner,duckdb_materializer,geometry_carriers,manifest_builder). - Contract clarification: no temporary feature flag is allowed; the rewritten key-first path becomes the default implementation.
- Contract clarification: discovery-driven schema extraction (column labels/descriptions/units) is required input to both UI payloads and column-selection validation for layers without explicit catalog
columns. - Contract clarification: baseline default exports must materialize only
subcatchmentsandchannelscarrier layers with carrier-grain row counts. - Contract clarification: temporal
event/yearlyexports must remain carrier-grain spatially and encode temporal selectors in wide measure column names (no long-format feature duplication). - Deliverable: maintainable and performant default export path with deterministic cardinality, deterministic naming, and verified small-watershed runtime targets.
WP-9: Service quality compliance closure and strictness coverage completion (planned 2026-03-28)
- Status: complete via
docs/work-packages/20260328_features_export_service_compliance_refactor/. - Scope: close QA "conditionally compliant" findings by extracting remaining service collaborators, removing dead wrappers, and adding missing strict-required branch coverage.
- Contract clarification: strict required-source behavior is shared across carrier and legacy paths through
discover_layer_sourcespolicy reuse; required source dependency/file/kind and unresolved required join key all fail explicitly withmaterialization_error. - Contract clarification: refactor does not change external submit/jobinfo/download contracts; service API behavior remains stable while orchestration internals are decomposed.
- Deliverable: reduced service orchestration complexity, closed medium QA findings, and preserved baseline run-path parity (
66/27onclogging-starch/disturbed9002-wbt-mofe).
WP-10: Standards-aligned artifact README.md generation and zip plumbing (planned 2026-03-29)
- Status: planned via
docs/work-packages/20260329_features_export_artifact_readme_metadata/. - Scope: generate deterministic, standards-aligned artifact
README.mdfrom resolved manifest/catalog/request metadata and package it into every features-export zip artifact. - Contract clarification:
manifest.jsonremains canonical for machine-readable consumers;README.mdis a human-readable derivative and must not diverge from manifest values. - Contract clarification: artifact bundles exclude profile files (
profile.yml, built-in profile files) while including generatedREADME.mdandmanifest.json. - Contract clarification: README generation must be deterministic and safe (no absolute paths/secrets), and cache-hit behavior must reuse the cached artifact README.
- Deliverable: end-to-end README generation helper + service packaging integration + regression coverage + documentation/work-package closure artifacts.
14.2 Dependency Order And Parallelism
- WP-1 depends on completed WP-0 Unitizer APIs.
- WP-2 depends on WP-1 outputs.
- WP-3 depends on WP-1; can proceed in parallel with late WP-2 testing once plan shape is stable.
- WP-4 depends on WP-2 and WP-3.
- WP-5 can begin after WP-1 contracts stabilize, then finalize against WP-4 endpoints.
- WP-7 depends on WP-4/WP-5 behavior and may revise portions of both.
- WP-8 depends on WP-7 outputs and supersedes WP-7 merge-path assumptions.
- WP-9 depends on WP-8 internals and closes remaining service-quality and strictness-coverage gaps.
- WP-10 depends on WP-9 baseline service contracts and updates artifact packaging/docs contracts.
- WP-6 final cutover validation follows WP-10.
14.3 Keep-It-Organized Rules
- Keep planner and validation logic pure and side-effect free.
- Keep file-system and geospatial I/O only inside exporter/service layers.
- Keep route/task files as adapters only; no business logic.
- Keep warning-code definitions centralized in
contracts.py. - Keep catalog/path resolution logic centralized in
catalog_loader.py; no duplicated path construction in writers. - Keep data-shaping joins/projections in DuckDB SQL paths; avoid pandas merge pipelines in export hot paths.
File-content cache identity amendment (implemented)
Service submissions MUST use SHA-256 for each existing regular file in the
resolved catalog dependency set. Keep mtime_ns in dependency manifests as
observability metadata. For a valid SHA-256 entry, the cache fingerprint uses
path/provenance, existence, size and hash, excluding mtime; metadata-only changes
therefore do not create another cache key. Explicit low-level none mode keeps
its legacy metadata identity. Missing entries and directory entries retain
explicit existing states; this amendment does not invent a content digest for
an absent file or a directory. Indirect vector/raster/directory dependency closure
is a separate open conformance obligation, not claimed by hashing the main path.
Hash reads use the verified bounded ordinary-file helper and its admitted-cache contract; metadata must agree around the hash observation. An invalid or missing hash never authorizes dropping metadata from a fingerprint. Existing source containment, parent-run roots, selectors, Unitizer settings, catalog/code versions and required-source validation remain unchanged.
Before publishing a newly materialized artifact into the reusable cache, re-read
the same catalog dependency snapshot and compare its fingerprint. A mismatch or
read failure MUST fail explicitly with FeaturesExportServiceError code
changed_source, status 409, and leave existing cache/artifact bindings intact.
Retain the candidate files and a manifest with dependency_verification.status
rejected (or error) and before/after fingerprint diagnostics; the RQ job fails
through its existing exception contract. On success, the manifest records
dependency_verification.status="verified" and the observed fingerprint.
Settle this verdict before writing any success manifest or README and before
packaging the ZIP, so artifact, bundle and job copies agree. On failure retain
artifact and job manifests with the same failure verdict; do not package or
publish a success bundle. This added manifest member is additive; old readers
may ignore it. Initial snapshot read failures also fail explicitly with
changed_source; never fall back to metadata-only identity after a hash error.
The recheck bounds the materialization window under existing producer behavior; it is not a multi-file filesystem transaction and cannot detect arbitrary change-and-restore entirely inside collection. A cache hit uses its initial verified dependency snapshot and immutable existing artifact; it does not recompute model output. Post-selection source changes are observed on the next submission. Job and published downloads remain historical artifact retrieval; this amendment does not require today's inputs to match a previously published artifact before downloading it.
New content fingerprints naturally miss prior metadata-only cache entries and rebuild on submission. Keep prior artifacts, manifests, cache entries, job URLs and publication bindings readable; do not migrate or manufacture historical hashes. Cache index and manifest schema versions remain compatible. Same-byte archive restoration can reuse a content-keyed artifact; actual changed source bytes require a new artifact even when size/mtime are restored.
Companion GeoPackage-to-FileGDB conversion MUST use the accepted producer's
provenance, never associate historical payload with today's source fingerprint.
Before conversion, require the source artifact's matching cache binding and a
verified content manifest. The binding must equal the corresponding current
GeoPackage request key (including Unitizer and version markers); its accepted
dependency fingerprint must equal the current companion dependency snapshot.
Resolve both plans from the same catalog; normalized requests must agree except
for format. Recheck the companion snapshot and request identity after conversion
before adding a reusable binding. Missing historical proof or a changed source
fails with changed_source 409, preserving previous published/cache bindings.
Use a distinct artifact candidate directory per conversion attempt; never
unlink or overwrite a previously accepted companion or the source GeoPackage.
Retain the rejected candidate and its own companion verification manifest when
conversion has produced files. Successful companion ZIPs add their own manifest
and README alongside the existing GDB tree; cache/result bindings reference
that companion manifest. Preserve the source GeoPackage manifest unchanged. Existing historical downloads remain available; do not infer
accepted hashes for legacy artifacts. Ordinary dual-format generation first
creates a newly verified GeoPackage and therefore remains supported. Verify
the companion before updating either profile publication registry entry, so
companion rejection preserves both prior published bindings. This ordering does
not promise a cross-file transaction for unrelated later I/O failures.
Profile publication derives request/dependency identity from the actual
artifact-matching cache binding, checking the requested format, instead of
collecting today's source snapshot. Publishing is selection of a completed
artifact, not certification against later model inputs. Reject missing or
incompatible bindings explicitly using the existing stale_publication contract.
The rationale for rejecting unsupported historical conversion is that old
manifests do not retain all Unitizer and request version inputs needed to derive
a different-format cache key honestly. Rebuilding through ordinary export is
the supported route; existing artifact retrieval does not need a rebuild.