Mitacs Globalink Research Internship · Ontario Tech Visual Computing Lab
A hyperspectral data investigation: SPLIB07 & ECOSTRESS
An honest, reproducible account of working with two public spectral libraries for contrastive spectrum-text learning — what the raw data actually looks like, the problems in it, the preprocessing pipeline built to clean it, the text enrichment used to make material names usable for language models, and what didn't work along the way.
Sources
Nothing here is our own measurement. Both libraries are redistributed with our own processing and metadata layered on top, never as a replacement for the original.
- USGS Spectral Library (splib07a)
U.S. government work, public domain. Redistributed here with our own derived metadata and citation.
- ECOSTRESS Spectral Library
JPL/Caltech, NASA contract. Reuse and redistribution permitted with attribution, which we provide on every page.
What the raw data actually looks like
Example filename from the raw ECOSTRESS archive: mineral.silicate.phyllosilicate.fine.tir.sepiolite_1.jhu.nicolet.spectrum.txt. Every sample's taxonomy, grain size, spectral range, sample ID, institution, and instrument is packed into the filename itself.
distributed as thousands of individual per-spectrum .txt files, not one table
roughly half the files are __MACOSX/._<name> resource-fork artifacts left over from how the archive was zipped on a Mac — no spectral data, pure overhead
the actual reflectance values and the metadata are split across separate files per sample
native resolution is far finer and inconsistent between SPLIB07 and ECOSTRESS, which is why a shared grid is necessary before any cross-library comparison
Preprocessing, with your own settings
Re-grid either library to a band count of your choosing and set a reflectance floor. This runs against the real processed arrays on the server, not a canned demo.
Making material names usable as text
Both libraries needed their material metadata turned into real text for contrastive spectrum-text learning, but for opposite reasons — and the fix differs accordingly.
SPLIB07 — synthetic enrichment (SPE)
SPLIB07 material names are sparse catalog labels (e.g. sepiolite_1) with nothing else attached. We generate a short, factual description per distinct material once — category, color, texture, composition — with gpt-4o-mini, and cache it. 633 / 633 materials enriched, 0 errors.
ECOSTRESS — real descriptions, extracted (no GPT)
ECOSTRESS already ships a real, curated description per individual sample from the original source metadata. The coarse material name isn't a safe key here — 420 of 752 distinct names cover samples with genuinely different write-ups — so this is keyed by filename, the actual unique identifier, and the text is extracted as-is rather than generated. 3,451 records, 100% with a non-empty description.
Known gap: the processed spectra array has no stored index back to which of these 3,451 records it came from, so these descriptions can't yet be paired 1:1 with a specific spectrum for training — only looked up by filename here. Closing that needs revisiting the original ECOSTRESS preprocessing step.
What didn't work
A research log is dishonest if it only shows what succeeded. This section documents approaches that were tried and abandoned, and why.
- TODO: name the approach
TODO: what was tried, why it seemed promising, and why it was abandoned.
Download
Every stage of the pipeline, available separately. Raw archives are redistributed as received from the original source with full attribution; everything downstream is our own processing.