Mitacs Globalink Research Internship · Ontario Tech Visual Computing Lab

A hyperspectral data investigation: SPLIB07 & ECOSTRESS

An honest, reproducible account of working with two public spectral libraries for contrastive spectrum-text learning — what the raw data actually looks like, the problems in it, the preprocessing pipeline built to clean it, the text enrichment used to make material names usable for language models, and what didn't work along the way.

Sources

Nothing here is our own measurement. Both libraries are redistributed with our own processing and metadata layered on top, never as a replacement for the original.

What the raw data actually looks like

Example filename from the raw ECOSTRESS archive: mineral.silicate.phyllosilicate.fine.tir.sepiolite_1.jhu.nicolet.spectrum.txt. Every sample's taxonomy, grain size, spectral range, sample ID, institution, and instrument is packed into the filename itself.

~45.7 MB
raw ECOSTRESS archive

distributed as thousands of individual per-spectrum .txt files, not one table

~50%
junk entries in the archive

roughly half the files are __MACOSX/._<name> resource-fork artifacts left over from how the archive was zipped on a Mac — no spectral data, pure overhead

2 file types / sample
spectrum.txt + ancillary.txt

the actual reflectance values and the metadata are split across separate files per sample

211 bands
the common grid we resample both libraries onto

native resolution is far finer and inconsistent between SPLIB07 and ECOSTRESS, which is why a shared grid is necessary before any cross-library comparison

Preprocessing, with your own settings

Re-grid either library to a band count of your choosing and set a reflectance floor. This runs against the real processed arrays on the server, not a canned demo.

Making material names usable as text

Both libraries needed their material metadata turned into real text for contrastive spectrum-text learning, but for opposite reasons — and the fix differs accordingly.

SPLIB07 — synthetic enrichment (SPE)

SPLIB07 material names are sparse catalog labels (e.g. sepiolite_1) with nothing else attached. We generate a short, factual description per distinct material once — category, color, texture, composition — with gpt-4o-mini, and cache it. 633 / 633 materials enriched, 0 errors.

ECOSTRESS — real descriptions, extracted (no GPT)

ECOSTRESS already ships a real, curated description per individual sample from the original source metadata. The coarse material name isn't a safe key here — 420 of 752 distinct names cover samples with genuinely different write-ups — so this is keyed by filename, the actual unique identifier, and the text is extracted as-is rather than generated. 3,451 records, 100% with a non-empty description.

Known gap: the processed spectra array has no stored index back to which of these 3,451 records it came from, so these descriptions can't yet be paired 1:1 with a specific spectrum for training — only looked up by filename here. Closing that needs revisiting the original ECOSTRESS preprocessing step.

What didn't work

A research log is dishonest if it only shows what succeeded. This section documents approaches that were tried and abandoned, and why.

Download

Every stage of the pipeline, available separately. Raw archives are redistributed as received from the original source with full attribution; everything downstream is our own processing.