fieldworkAGRICULTURAL INTELLIGENCE

METHODS / A REPRODUCIBLE PUBLIC-DATA PROJECT

From machine records
to a defensible field trial.

First, inspect what was aligned, flagged and withheld. Then see how plot-scale variation determines the replication needed for a small yield difference.

52,827original yield observations
2,724unique QA-flagged records
2 harvestsone pinned public source

Measured machinery findings. Prospective power scenarios. No client trial, calibrated sensor correction or causal yield gain is claimed.

01 / MACHINERY DATA

Align the frame.
Keep the audit trail.

2018 harvest positions are in WGS84 degrees; application points are in UTM 12N meters. We reproject the harvest into the application frame. The measurements remain intact.

Same coordinate frame, extent and yield scale within each harvest.

ORIGINAL RECORDS 31,276
2018 original yield records; quality flags shown in orange
QA RETAINED 30,133
2018 yield records passing the balanced quality rules
QUALITY SCREENING1,143

unique records flagged

ACCEPTED SPATIAL MATCH50.8%

15,302 of 30,133 QA-retained records

JOIN RULE / DATA LIMIT

Nearest application point within 12 m. Coordinate reprojection does not establish GNSS accuracy.

What was found, and what we did.

Diagnostic20182016Action
Coordinate-system mismatchDegrees vs metersSame source projectionTransform 2018 to EPSG:32612; preserve the unspecified 2016 datum.
Duplicate x / y / yield5197Keep first in the analysis view; preserve the full ledger.
Non-positive yield or flow01,146Flag and exclude; no replacement values.
Speed below 1.5 mph342Not suppliedScreen 2018; no invented 2016 speed.
Moisture outside 8–22%24782Flag outside the stated range.
Non-positive swath0833Flag and exclude; source units retained.
Yield percentile tails6261,501Screen outside central 98% of positive values.
Conflicting polygon ratesNot applicable592 raw casesWithhold joins when overlap rate spread exceeds 5 source lb/ac.

Quality reasons overlap. These screening choices can flag legitimate extremes; they are not all proven machine errors. All original rows remain inspectable.

COVERAGE / 2018

A wider radius does not repair the gap.

8 m matches 47.0% of QA-retained yield; 12 m matches 50.8%; 20 m matches 51.3%. Only 162 records are added from 12 to 20 m.

Proximity alone does not prove time alignment or randomized treatment assignment.

ACCURACY / UNRESOLVED

Small separation is not independent control.

13,136 QA-retained harvest points are within 1 mm of an application coordinate. Numerical coincidence in source positions is not evidence of sub-millimeter GNSS accuracy.

Sensor delay, GPS drift, antenna offsets and yield calibration need supporting instrument evidence before a physical correction can be claimed.

The 2018 GeoJSON retains original WGS84 positions. No WGS84 layer is exported for 2016 because its datum is unspecified.

02 / FIELD-TRIAL REPLICATION

Count independent plots.
Start with the variance.

Two treatments, randomized within independent paired blocks. Each block contains one A plot and one B plot. GPS records within those plots are subsamples.

Two-sided alpha = 0.05. Exact paired-t power under normal, independent block differences. These are scenario inputs, not measured client variance.

COMPLETE PAIRED BLOCKS REQUIRED

52

104 total plots / 52 per treatment

80.78% calculated power

Before plot or pair losses. Validate plot-scale variance and the planned design before choosing a field layout.

Power for a 2 bu/ac difference: 15, 52 or 199 paired blocks reach 80% as difference SD increases from 2.5 to 5 to 10 bu/ac
FIXED EFFECT: 2 BU/AC / 80% TARGET LINE / TWO PLOTS PER BLOCK

A BETTER VARIANCE ESTIMATE

Use pilot plot means.
Test the proposed design.

Estimate the SD of paired treatment differences from comparable replicated trials or a pilot. If the two plot SDs equal σ and their within-block correlation is ρ, then SD(D) = σ√[2(1−ρ)].

The historical 20 m-cell residual SD is a proxy, not a validated estimate of independent trial-plot variation. Multi-farm, multi-season and spatially correlated designs need their own analysis.

IMPLEMENTATION / VERIFICATION

Plan for operations.
Check calculation behavior.

52 complete pairs with an assumed 10% pair loss means scheduling 58 blocks / 116 plots for expected retention. This is not a guarantee. Machine width, plot size, buffers and harvest calibration still need a protocol.

20,000 simulations per validation scenario agree with exact power within four Monte Carlo standard errors. Simulated zero-effect rejection was 4.88% against nominal 5%.

SOURCES / REPRODUCIBILITY

Evidence you can take with you.

The 8-page report contains the full handling rules, match sensitivity, assumptions, replication table and references. The project includes source hashes, audit rows, scripts and original frozen files.

  1. OFPEDATA / Montana State University / MIT / frozen source ↗
  2. pyproj / coordinate transformer ↗
  3. University of Minnesota Extension / on-farm research design ↗
  4. SciPy / noncentral-t power calculation ↗