Fix CV7000 sparse acquisition series and metadata handling#4450
Fix CV7000 sparse acquisition series and metadata handling#4450Darlokt wants to merge 45 commits into
Conversation
Build CV7000 series from acquired well/field pairs with readable TIFF planes instead of expanding every acquired well to the global maximum field index. Update SPW metadata to emit Image, WellSample, and PlateAcquisition references only for acquired fields, using compact per-well sample indexes. Set PlateAcquisition MaximumFieldCount from the maximum acquired field count per well, and make missing-plane duplication opt-in so missing planes return fill pixels by default.
Populate channel and objective metadata from representative real planes instead of assuming the first plane index for each channel is present. Match channels by raw channel plus timeline/action, falling back to raw channel only when unambiguous. Skip plane position metadata for missing planes so positions are only written for real backed planes.
Populate WellSample positions only from TIFF-backed planes, avoiding metadata-only missing planes in sparse acquisitions. Set per-series physical pixel sizes from the first backed channel with resolvable channel metadata instead of requiring channel 0 to be present.
Preserve per-plane position metadata for mapped MLF IMG records even when the referenced TIFF is missing, so dense logical planes with fill pixels keep their own parsed X/Y/Z coordinates. When resolving representative planes for channel metadata, prefer TIFF-backed planes but fall back to metadata-only planes for the same channel if no backed plane is available.
Recent sparse CV7000 fixes changed per-series metadata lookup to match channels by raw channel plus timeline/action. That preserved channel ordering for sparse acquisitions, but exposed a parser ordering issue: timeline/action channel instances are created while parsing Timelapse, before ChannelList is parsed. As a result, exact matches for later timelines resolved to channel instances without objective, exposure, color, or light source metadata. Propagate ChannelList settings to every channel instance with the same raw channel index, and add light source references to all matching instances without duplicates. Keep exact timeline/action lookup first, but let raw-channel fallback prefer a populated channel if exact lookup is unavailable. Deep-copy light source references when duplicating channels so later updates do not share mutable state accidentally. This keeps the sparse acquired-field series allocation intact while restoring optical metadata for all timeline/action channel occurrences.
|
Thanks, @Darlokt. Before we can consider including this, we do need:
See also https://ome-contributing.readthedocs.io/en/latest/code-contributions.html#procedure-for-accepting-code-contributions. Please let us know when both the CLA and test data have been submitted. |
|
Thanks @melissalinkert, the CLA is signed! |
|
I haven't been able to find a record of the CLA being submitted. Can you please double-check that it was submitted as instructed at https://ome-contributing.readthedocs.io/en/latest/cla.html? We really do need a test case of some sort that both demonstrates why these changes are necessary and can be included in our test data repository in the future. Would you be able to replace the images in your test datasets with something artificial (see https://bio-formats.readthedocs.io/en/latest/developers/generating-test-images.html), and submit the modified dataset? |
|
I double checked, it was send to the correct address, but I can send it again? I can try to make a minimal synthetic test set, but I cant say its going to be correct, as I don't have the microscope to double check. It mostly boils down to .mrf vs .mlf mismatch and when compensating for that the downstream things that break. I can try to write a minimal test set based in this? |
Use Yokogawa MES bts:SliceLength as the physical Z spacing in micrometers, based on the CV7000 metadata and observed MLF Z positions. Parse the value from acquire actions and attach it to the mapped timeline/action channel metadata. Set PixelsPhysicalSizeZ when all represented channels in a series agree on the same slice length. Leave it unset for missing or conflicting slice length metadata, and keep per-plane Z positions written in the existing reference-frame units.
Use Yokogawa MES light-source links as channel excitation metadata and treat Acquisition filter labels as detection filters. Populate emission wavelengths for band-pass filters, create OME Filter and LightPath emission filter references, and add support for BP, LP, and SP filter ranges. Create Detector metadata from Yokogawa camera numbers and link per-channel DetectorSettings with binning. Preserve ambiguous gain, lamp, filter, and camera provenance as OriginalMetadata while leaving brightfield lamp channels unmapped to core OME light sources.
Split CV7000 initialization into explicit dataset path, sidecar parsing, series layout, plane lookup, and metadata population phases. Track series timing from measurement planes and populate per-image acquisition dates and plane delta times from Yokogawa timestamps. Preserve Yokogawa sidecar provenance as OriginalMetadata, including MRF, MES, WPP, channel action settings, measurement summaries, and chunked raw XML with SHA-256 digests. Add structured objective, acquisition mode, contrast method, detector, binning, light path, and filter metadata population on top of the refactored reader layout.
Normalize Yokogawa decimal-comma numeric values before parsing CV7000 sidecar floats, including plane positions, pixel sizes, light source settings, exposure times, filter wavelengths, and Z spacing. Parse Andor gain parameters into DetectorSettings.Gain while retaining the original Andor parameter fields as provenance.
e1e4d5d to
a22966b
Compare
After checking DataTools.parseDouble does everything needed, the wrapper was unnecessary and redundant.
Preserve `PostProcess.ppf` as raw CV7000 sidecar XML and parse lightweight PostProcess summary/action metadata while leaving verbose logs unexpanded. Keep OME `PhysicalSizeZ` limited to shared positive channel Z spacing, with per-channel Z offsets represented through per-plane `PositionZ` to account for per channel differing z spacing.
Restrict CV7000 dataset sidecar discovery to the declared acquisition files plus known Yokogawa sidecars, including PostProcess.ppf, OTF_crosstalk_parameter.xml, and OTF_geometry_parameter.xml. Avoid adding unrelated XML files such as generated OME-XML metadata while preserving raw metadata for recognized sidecars and treating both .tif and .tiff as pixel files.
Parse OTF_crosstalk_parameter.xml from the .wpi directory to populate exact emission filter ranges, transmittance, and single-dichroic light path references for matching CV7000 channels. Preserve crosstalk filter, dichroic, fluorophore, and intensity values as Yokogawa original metadata while distinguishing filters by camera number. Classify CV7000 companion files by dataset role so no-pixels file lists exclude TIFF planes consistently and known Yokogawa sidecars are matched by resolved dataset paths instead of recursive basename discovery. Preserve raw XML only for explicit sidecar roles, leaving MeasurementData.mlf and parsed PostProcess.ppf out of the raw XML dump while retaining their structured original metadata.
|
Thanks @Darlokt, I confirm we have received your CLA (it just took a while to appear). Yes, please do write relevant tests that cover all of the changes in this pull request. This is a large set of changes to an established reader, so we need concrete tests to understand what wasn't working previously. If looks like you're still actively working on this, so I would ask that you please convert this pull request to a |
Set CV7000 plate well origins to zero and store MLF field X/Y values as micrometer well-sample offsets instead of reference-frame coordinates. Treat those positions as well-center-relative field offsets because the sidecar ranges fit inside the WPP well geometry only when interpreted as micrometers. Stop writing per-plane X/Y positions from the same field offsets, leaving them on WellSample where they describe the acquisition field. Store MLF Z positions as micrometer AF-relative stack coordinates rather than generic reference-frame positions.
Store CV7000 MLF Z values as per-plane reference-frame coordinates instead of micrometer stage positions. Keep calibrated Z spacing on Pixels.PhysicalSizeZ when available, since MLF Z is AF-relative and not proven absolute stage Z.
Set OME plate row and column naming conventions for CV7000 datasets, record well-sample timepoints alongside image acquisition dates, and replace generic channel names with display labels derived from brightfield mode, target, and acquisition metadata. Preserve the original action/channel/camera identifiers as channel provenance metadata so the raw CV7000 source mapping remains available
Replace chunked global raw-XML metadata with structured CV7000 MapAnnotations plus raw sidecar FileAnnotations. Summarize sidecars with role, byte length, and SHA-256, link scoped annotations to plate, instrument, image, and channel objects, and keep large XML payloads attached as binary annotations instead of split base64 metadata.
Split CV7000Reader’s transient parsing, layout, and original-metadata bookkeeping into dedicated reader-local model/helper classes. Centralize dataset paths, raw sidecar state, acquired-well/field indexes, representative-plane lookup, and shared Yokogawa parsing utilities so metadata population can reuse indexed state instead of rescanning planes or duplicating string-parsing logic.
Record non-laser CV7000 illumination in OME metadata instead of only emitting laser-backed channel settings. Map brightfield channels to lamp light sources, write lamp entries as halogen filaments, and keep laser excitation wavelengths limited to true laser sources. Also improve related instrument metadata by normalizing Andor camera handling and preserving Yokogawa filter-wheel labels on generated filters.
Map CV7000 white-light sources to OME FilamentType=Other instead of Halogen. The manual describes a halogen lamp, but MES sidecars expose only generic Type="Lamp" and installed hardware may differ, so the reader should avoid asserting a more specific filament type than the source metadata supports.
Add a cv7000.preserve_raw_sidecars metadata option and keep raw CV7000 sidecars attached as file annotations unless callers explicitly disable it. Also clean up parser helper usage, guard sidecar handling against null entries to avoid a possible null pointer, and remove the unused well-acquisition helper while keeping the default metadata output unchanged.
Parse the CV7000 OTF geometry sidecar and use it to populate objective records, including inferred nominal magnification, lens NA, and immersion when channel metadata alone is incomplete. Also enrich promoted optics metadata with laser model/type, lamp power, channel illumination/acquisition/contrast fields based on information form there and the CV7000 user manual, and preserve matched OTF geometry affine rows in original metadata for auditability.
Map MRF target system values onto the OME microscope model with Yokogawa as the manufacturer. Preserve the cleaned operator name as an Experimenter and link it from each image, while continuing to warn on non-CV7000 target systems.
Track duplicate-plane selection while building CV7000 plane lookups and publish per-image provenance summaries in Yokogawa original metadata. Annotate TIFF-backed, metadata-only, duplicated, filled, and missing-MLF planes so sparse acquisitions and ignored duplicate candidates remain auditable from the reader output.
Parse observed CV7000 objective names into magnification, immersion, phase, and long-working-distance tokens, then map known combinations to nominal magnification and lens NA when populating OME metadata. Refactor sidecar parsing into typed result objects, tighten helper visibility, and stream sidecar SHA-256/length summaries instead of buffering full files.
Reorganize `CV7000Reader` into clearer functional modules, extract shared channel-mapping and SAX parsing helpers, and centralize attribute cleanup through `YokogawaParsing`. Reset all parsed reader state when reopening datasets on a reused reader so CV7000 sidecar, channel, and OTF metadata do not leak across opens.
Add a `cv7000.infer_objective_lens_na` metadata option and default it to off. Only write `Objective.LensNA` and the `MappedLensNA` Yokogawa metadata when that option is explicitly enabled, since the CV7000 objective mapping is inferred rather than directly recorded in the acquisition metadata.
Replace CV7000’s custom plane index math with `FormatTools.getIndex` and `getZCTCoords` using an explicit `XYZCT` dimension order. This aligns reverse plane lookup, duplicate-plane fallback, and per-plane metadata with the reader’s logical Z/C/T layout, follows OME-Zarr’s preferred axis order, and fails fast if a Yokogawa channel cannot be mapped into the compact reader channel set.
Use action-aware channel indexing only when every acquired plane has a complete MES-derived action/channel match. Fall back to raw MLF channel indexes when channel sidecar metadata is missing, lacks action assignments, or does not map all acquired planes, and record the fallback mode and reason in Yokogawa provenance metadata.
Treat CV7000 `bts:Power` values as light-source attenuation rather than physical laser or lamp power. Store valid values on `Channel.LightSourceSettings.Attenuation`, preserve them in original metadata, and stop populating OME Laser/Filament power fields from these percent-like sidecar values. Also mark CV7000 lasers as `SolidState` based on the Yokogawa instrument documentation.
Keep CV7000 channel correction file metadata relative to the WPI directory while storing a resolved absolute path for used-file detection. This avoids leaking local absolute paths into metadata without losing the ability to report existing correction sidecars.
Mark preserved CV7000 XML sidecar BinaryFile BinData as little-endian when storing raw bytes, avoiding missing required BinData endianness metadata in generated OME annotations.
|
Hej @melissalinkert , sorry, for the delay. While creating the synthetic dataset fixtures I realized that the old reader was not maintaining all associated metadata, so beyond fixing the reader to handle sparse acquisitions, I also rewrote the internal model to preserve more Yokogawa metadata and make the reader easier to maintain, hopefully reaching archival quality, so the old original data can be retired, with everything associated to it being preserved in the new reader. Here a more detailed overview of the changes and their mappings. As part of responsible disclosure, an AI coding Agent helped with the writing of this overview/review and as a secondary code reviewer throughout the rewrite. New Reader OverviewThe old reader could open the basic file family, but most of the reader behavior lived inside one large The new reader keeps the same reader target and file family, but changes the internal model. The central behavioral change is conservative pixel exposure. Executive SummaryCompared with the old reader, the new reader:
Problems in the old readerDense field expansionThe old reader collected acquired wells and the largest observed field index. It then allocated That is only correct for dense acquisitions where every acquired well has every field slot from The result was empty Bio-Formats series and Metadata-only records could shape pixel layout
That meant a record with no readable pixels could affect:
The new reader keeps metadata-only records available for metadata and provenance, but TIFF-backed acquired fields decide series allocation. Channel metadata lookup was too broadThe old reader used raw Yokogawa channel number as the primary channel lookup key. That was fragile when the same raw channel was acquired in multiple timelines or actions with different optical settings. It could attach channel metadata from the wrong action to a logical plane. The old metadata loop also used Missing-plane duplication was enabled by defaultThe old default for That behavior invents pixel data that was not acquired. Sidecar metadata was underusedThe old reader parsed the main plate, measurement, and measurement-setting sidecars, but much of the available CV7000 sidecar metadata was not promoted into standard OME structures. Missing or weakly represented metadata included:
The old output either omitted these values, flattened them into generic metadata, or made them difficult to audit from exported OME-XML. Sparse metadata gapsBecause channel metadata depended on fixed logical plane positions, otherwise available metadata could be skipped for sparse acquisitions. Specific gaps included:
New Reader ArchitectureThe old reader performed parsing, layout, reverse lookup construction, and OME metadata population as one interleaved flow. flowchart TD
A["Open .wpi"] --> B["Parse plate"]
B --> C["Parse MLF records"]
C --> D["Parse MRF and MES"]
D --> E["Scan planes"]
E --> F["Collect acquired wells"]
E --> G["Find max field"]
F --> H["Allocate dense series"]
G --> H
H --> I["Build lookup"]
I --> J["Populate OME inline"]
The key problem is the dense allocation step. the old reader mixed "this well was acquired" with "all field slots up to the global maximum were acquired". The new reader separates parsing, layout, lookup construction, metadata population, and provenance emission into staged responsibilities. flowchart TD
A["Open .wpi"] --> B["Reset state"]
B --> C["Resolve dataset paths"]
C --> D["Parse WPI and MLF"]
C --> E["Parse optional sidecars"]
D --> F["Build raw model"]
E --> F
F --> G["Normalize channels"]
D --> H["Build acquired-field layout"]
G --> H
H --> I["Initialize CoreMetadata"]
I --> J["Populate plane lookup"]
J --> K["Populate OME metadata"]
F --> K
K --> L["Emit Yokogawa provenance"]
The staged model gives each phase one main decision:
The reader source now marks the implementation into modules:
The important design change is ownership. Pixel and Series BehaviorThe new reader changes series allocation from dense inferred fields to exact acquired fields. flowchart TD
A["MLF plane records"] --> B{"Readable TIFF?"}
B -->|yes| C["Add exact well/field"]
B -->|no| D["Keep metadata-only record"]
C --> E["Sort row/column/field"]
E --> F["Create acquired-field series"]
D --> G{"Same field has TIFF?"}
G -->|yes| H["Can contribute metadata"]
G -->|no| I["No Image or WellSample"]
F --> J["Build logical plane lookup"]
H --> J
The difference is visible in the high-level behavior:
Reverse Plane LookupThe reverse lookup remains the bridge between Bio-Formats logical plane index and Yokogawa plane record. The rewrite makes the assignment rules explicit. flowchart TD
A["Plane record"] --> B{"Acquired field?"}
B -->|no| C["Skip lookup"]
B -->|yes| D["Map channel"]
D --> E["Compute XYZCT index"]
E --> F{"Slot empty?"}
F -->|yes| G["Select record"]
F -->|no| H{"Candidate has TIFF?"}
H -->|yes| I["Replace metadata-only or ignore duplicate TIFF"]
H -->|no| J["Ignore metadata-only duplicate"]
I --> K["Record duplicate candidate"]
J --> K
G --> L["Index representative plane"]
K --> L
A metadata-only record can no longer permanently occupy a logical slot when a later TIFF-backed record maps to the same slot. If a duplicate TIFF-backed record is ignored, it is still recorded as a duplicate candidate and reported as provenance. Missing Plane ReadsThe public option still exists The default changes from When duplication is explicitly enabled, the reader tries to duplicate a TIFF-backed plane from logical Z 0, the same channel, and timepoint 0. The duplication path verifies that the candidate has a real file before recursively reading it. flowchart TD
A["openBytes"] --> B["Fill buffer"]
B --> C{"Requested plane has TIFF?"}
C -->|yes| D["Read TIFF"]
C -->|no| E{"Duplication enabled?"}
E -->|no| F["Return fill"]
E -->|yes| G["Find Z0/C/T0 source"]
G --> H{"Source has TIFF?"}
H -->|yes| I["Read source plane"]
H -->|no| F
This makes data integrity the default while preserving the old visualization escape hatch for users who choose it. Channel Mapping and Representative PlanesThe old reader had two separate channel problems:
The new reader centralizes this behavior in flowchart TD
A["MRF channels"] --> B["Base channel objects"]
C["MES actions"] --> D["Timeline/action channel copies"]
B --> E["Normalized channel list"]
D --> E
E --> F["Index timeline/action/raw channel"]
G["MLF records"] --> H{"Complete action mapping?"}
F --> H
H -->|yes| I["ACTION_MAPPED channels"]
H -->|no| J["RAW_MLF fallback"]
I --> K["Compact channel axis"]
J --> K
K --> L["Representative plane per series/channel"]
L --> M["OME channel metadata"]
The lookup order for OME channel metadata is:
The layout stage uses action-mapped logical channel indexes only when every MLF record in an acquired field can be mapped. That avoids a partially remapped channel axis, where some planes would use action-aware channel order and others would use raw MLF numbering. Representative planes are indexed during reverse lookup. If no representative plane or matching channel metadata can be found for an OME channel, that channel metadata slot is left unset instead of borrowing metadata from an unrelated channel. Metadata Added by the RewriteOME population now runs after sidecars, series layout, core metadata, and the reverse plane lookup are ready. flowchart TD
A["CoreMetadata ready"] --> B["Populate Pixels"]
B --> C["Populate Plate and Wells"]
C --> D{"MINIMUM metadata?"}
D -->|yes| E["Stop after sparse HCS layout"]
D -->|no| F["Populate Instrument"]
F --> G["Populate Experimenter"]
G --> H["Populate Image, Channel, Plane"]
H --> I["Emit Yokogawa annotations"]
The The full metadata path creates shared instrument IDs and indexes before populating each series. Those indexes let channels link to the correct light source, detector, filter, dichroic, and objective metadata. Source-to-OME Mapping
Plate, Image, Pixels, and PlanesThe new reader keeps the reader in the HCS domain and exposes one OME Plate and well metadata changes include:
Each acquired well/field pair becomes one OME Logical raster calculations now use Bio-Formats dimension-order helpers:
Plane and timing metadata changes include:
The X/Y change is intentional. CV7000 sidecars mix coordinate systems. Instrument, Objectives, Detectors, and Light SourcesThe new reader creates an Instrument metadata can include:
The old reader only created an instrument when laser light sources were present. Objective metadata now comes from MES channel metadata and OTF geometry affine rows keyed by objective ID. The option Detector metadata is promoted from camera metadata. The MRF operator name is emitted through the separate Channels, Filters, Light Paths, and ExposureChannel, light path, filter, dichroic, detector, and exposure metadata now share one representative-plane lookup path. The new reader can populate:
The semantic change is that excitation now comes from the light source linked to the channel. Most outputs remain conditional:
If the representative plane is missing, or if channel lookup cannot find a matching channel, the reader skips that OME channel metadata slot rather than copying unrelated metadata. Yokogawa Provenance and Raw SidecarsThe new reader introduces grouped Yokogawa metadata annotations. Instead of emitting one large flat original-metadata set, related values are grouped by scope and emitted as OME Examples of scopes include:
flowchart TD
A["Parsed sidecar fields"] --> B["Annotation group"]
C["Plane provenance"] --> B
D["Sidecar summary"] --> B
B --> E["MapAnnotation"]
E --> F{"Best scope"}
F -->|plate/general| G["Plate ref"]
F -->|instrument/OTF| H["Instrument ref"]
F -->|image/action| I["Image ref"]
F -->|channel| J["Channel ref"]
This keeps standard OME metadata focused on interoperable fields while keeping Yokogawa-specific source fields and reader decisions available in exported OME-XML. Plane ProvenanceOutside Each summary includes:
When the field can be resolved, the summary also includes:
When a plane is filled, duplicated, or a duplicate logical record is ignored or replaced, the summary can include repeated Duplicate-candidate anomalies use a common
Plane provenance is informational. It does not alter Raw Sidecar PreservationThe new reader adds The preservation policy is selective:
All classified non-pixel sidecars can still receive summary provenance with role, byte length, and SHA-256. Raw preservation controls embedded payloads, not whether sidecars are summarized. Annotation ShapeThe provenance output uses standard OME annotation structures:
Plane anomaly rows are encoded as compact semicolon-delimited strings in Reader Options and Used FilesThe reader now exposes three CV7000-specific options:
The pixel-affecting default is conservative: missing pixels are not invented by default. The new reader classifies known CV7000 files by role instead of relying on suffix checks alone.
The reader can now:
Known dataset roles include:
Internal StructureThe rewrite adds reader-local structures that make the data flow explicit:
The model boundaries matter more than the helper class names. The reader now has a clear split between:
That separation makes the sparse acquisition behavior auditable. Behavior Comparison Summary
Overall EffectThe new reader makes the CV7000 reader more conservative about pixels and more complete about metadata. It no longer invents empty field series from nominal or maximum field counts. The rewrite also makes channel metadata safer. At the OME layer, the reader now exposes richer plate, timing, channel, instrument, objective, detector, filter, dichroic, light source, physical size, and provenance metadata. The result is a reader that better represents sparse CV7000 acquisitions as they were actually acquired, while making sidecar-derived metadata and reader decisions auditable from exported OME-XML, reaching by my observation full preservation and archival quality, for full conversion fo exsting data without loosing metadata. |
|
As part of this I have made minimal synthetic dataset-fixtures to check for regressions, with the associated As real data I found available S-BIAD1013 available open at the BioImage Archive which could be included for real data testing. The other studies I found do not publish the metadata files. S-BIAD1013 passed the reader tests fully and was fully parsed. Beyond that, I would love some guidance on the NA inference path. In a global view the metadata is still quite incomplete, no disk data, back projected pinhole etc. is available etc. but that is a problem of the metadata format, as no more information can be stably inferred, at least from my point of view. |
A single submission to https://zenodo.org/communities/bio-formats is fine. |
|
@melissalinkert Perfect! The datasets are under https://zenodo.org/records/21224717. |
Store CV7000 dataset paths and measurement summaries in memo-friendly value objects instead of retaining Location and SAX handler instances. Make reader-local helper models accessible to Kryo with public classes, fields, and no-arg constructors, avoiding Java module-access failures when saving bioformats2raw memo files on Java versions newer than 11.
|
ed3fbc6 fixes memoization and local helper model access for Kryo in Java > 11. |
Assign stable logical channel slots across action and timeline acquisitions, qualify overlapping timeline channels, and preserve raw-channel fallback behavior. Match repeated MRF channel definitions by occurrence, mark ambiguous metadata safely, and improve duplicate-plane reuse by selecting TIFF-backed planes from the same logical channel.
Apply consistent Yokogawa channel bit depths to each series, validate values against storage width, and retain TIFF-derived fallbacks when channel metadata is missing, ambiguous, invalid, or conflicting.
Log missing sidecar metadata and unavailable exact acquisition matches, while deduplicating fallback warnings per raw channel. Guard channel lookup against null planes and channel lists, and reset warning state when reindexing metadata.
Add configurable Zlib compression for preserved raw sidecars, recording the compression type and encoded payload length while retaining an option to store the original bytes.
Include measurement data and post-process XML roles when identifying CV7000 metadata files, to allow embedding. With compression options, this seems now like a possible option.
Add namespace-independent Yokogawa XML parsing helpers and use them consistently across CV7000 sidecar handlers. Parse and summarize MLF and post-processing provenance, link record annotations to matching images, and limit emitted detail records through a configurable metadata option.
Track acquisition and shading-correction TIFFs separately instead of treating every TIFF beside the WPI as a dataset plane. Only accept readable TIFFs referenced by MLF IMG records as acquisition files, while retaining referenced correction TIFFs as non-plane dependencies.
Recover uniquely matching readable TIFF corrections when expected files are absent, while rejecting ambiguous matches and warning about their sources.
|
Hej @melissalinkert, Since my previous message, I have added nine follow-up commits. These changes refine the channel model, make metadata fallback behavior more explicit, extend sidecar provenance, and tighten the distinction between acquisition pixels and other TIFF dependencies. Channel mapping and bit depthThe reader now assigns stable logical channel slots across actions and timelines. A channel is only timeline-qualified when the same acquisition coordinate would otherwise collide across timelines; non-colliding channels retain their primary slots. This avoids both merging distinct acquisitions and unnecessarily expanding the channel axis. Repeated MRF channel definitions are matched by occurrence rather than always selecting the first raw-channel match. Ambiguous definitions are marked as such and are not used to populate potentially incorrect OME channel metadata. Missing-plane duplication, when explicitly enabled, also now searches for a TIFF-backed plane in the same logical channel instead of relying on an earlier fixed position. The reader additionally prefers Yokogawa Channel lookup fallbacks are now visible in the log. When no exact timeline/action/raw-channel match is available and the reader has to use a broader raw-channel definition, it emits one deduplicated warning per acquisition/channel key. This preserves the existing safe fallback while making degraded metadata resolution auditable. Sidecar preservation and structured provenanceRaw sidecar preservation now supports Zlib compression through the new With the savings from the zlib compression, it now makes sense to also embed the bigger files, so now The XML parsing code has been centralized around namespace-independent element and attribute helpers and is now used consistently by the Yokogawa sidecar handlers. On top of the raw preservation, the reader exposes more structured provenance:
Detailed MLF and PPF records are bounded by the new TIFF and shading-correction discoveryTIFF discovery is now reference-driven. The reader tracks acquisition TIFFs referenced by MLF If a referenced shading-correction file is missing under its exact metadata name, the reader searches the expected directory for readable TIFFs whose filename stem contains the expected stem. This was implemented, as in some cases they can ave prefixes, I am unsure where they come from, but now this possibility is handled. It accepts the fallback only when the match is unique, updates the resolved dependency path, and warns rather than guessing when multiple candidates match. This handles CV7000 packages in which correction files have gained an additional filename pre/suffixes while keeping resolution deterministic. CleanupThe follow-up series also removes obsolete plate-description and measurement-summary accessors that were no longer used after the reader restructuring, and changes internal used-file collections to insertion-ordered sets to avoid duplicate entries while retaining deterministic ordering. SummaryOverall, these commits do not change the general previous behavior of the reader but they make its channel axis more robust for repeated and multi-timeline acquisitions, implement/improve metadata based significant-bit reporting, preserve and summarize more of the original Yokogawa audit trail, and tighten file discovery that only metadata-referenced TIFFs are considered part of the dataset. In line with this I have also extended/improved the synthetic test cases, adding them under a new version |
Hej everyone,
I was doing some conversions of some CV7000 data and found some weird behavior in the metadata construction etc. related to sparse acquisitions and timeline/action-specific channel metadata.
The reader previously assumed dense well/field layouts and often derived metadata from fixed plane positions, which broke when datasets had missing TIFFs, skipped fields, or timeline-specific channel instances.
The changes make series allocation, plane lookup, SPW metadata, channel metadata, physical pixel sizes, and missing-plane handling derive from actually acquired/readable planes where appropriate, while preserving parsed metadata-only plane records for fallbacks.
Problems Found
Fixes
I had a bit of a problem with plane duplication being enabled by default. According to 5cd0ea5 this was added a a better visualization option, but left enabled by default. From where I stand, the default behavior should never create false data.
I did a small top level search throughout and also found similar behavior in the InCell reader. If you want I can also revert this behavior there to disabled by default.
After the fixes, on my test data, the conversion now creates the proper number of files associated with the actual number of planes captured as described by the Yokogawa metadata with correctly populated metadata.
All the best!
Kilian