upxo.pxtalops.grain_splitting_3d module

Geodesic within-grain splitting for 3D labeled grain structures.

Subdivides each grain in a 3D labeled voxel image (LGI) into one or more connected sub-regions via multi-source BFS (geodesic nearest-seed assignment), seeded from farthest-point-sampled voxels within the grain itself. Standalone and framework-agnostic (pure numpy + stdlib) – takes and returns plain label arrays and dicts, with no dependency on any specific grain-structure class, so it is reusable across UPXO’s various 3D pxtal/grain-structure modules (originally built for, and still used by, the fm_steel_3d PAG/packet pipeline’s “Technique B”, but the algorithm itself is domain-agnostic).

Guarantees, by construction:
  • Every output sub-region is a single connected component (a voxel can only be reached by travelling through voxels that belong to the same original grain).

  • Sub-regions exactly partition their parent grain’s voxels (none lost, none duplicated).

  • No dependency on inter-grain adjacency/connectivity – splitting one grain never competes with any other grain for resources, unlike clustering-based approaches to forming multi-voxel-region groups.

This is one splitting strategy among several possible ones – geodesic nearest-seed assignment from spread-out (farthest-point-sampled) seeds gives reasonably balanced sub-region sizes (an irregular grain shape or unlucky seed placement can still skew sizes), not an exact equal-volume partition. A different algorithm (e.g. iteratively-rebalanced/weighted seed placement, or a power-diagram-style approach) would be needed for a guaranteed equal-size split; that is a natural direction for future development in this module without changing its calling contract. size_balance_metrics (below) exists precisely to quantify that gap – how far a given split landed from an equal one – so it can be tracked as splitting strategies evolve.

upxo.pxtalops.grain_splitting_3d.split_grains_geodesic(lgi: numpy.ndarray, size_distribution: Dict[str, Sequence[float]], min_voxels_to_split: int, connectivity: int = 6, random_seed: int | None = None) → Tuple[numpy.ndarray, Dict[int, int], Dict[int, List[int]]][source]

Subdivide every grain in lgi into 1 or more sub-regions.

Grains with fewer than min_voxels_to_split voxels are left whole (single sub-region == the whole grain; this is the correct, intended outcome for a genuinely small grain, not a degenerate/error case). Grains at or above the threshold get a sub-region count sampled from size_distribution, then their voxels are partitioned into that many sub-regions via multi-source BFS (geodesic nearest-seed assignment) seeded from farthest-point-sampled voxels within the grain – guaranteed connected sub-regions by construction, since a voxel can only be reached by travelling through voxels that belong to the same original grain.

Parameters:
  • lgi (np.ndarray) – 3D labeled grain image. 0 = background, 1..N = grain ids. Every nonzero label is assumed to already be a single connected component at the given connectivity (true for any cc3d.connected_components or Voronoi-argmin output) – this is what guarantees every voxel in a grain ends up assigned to some sub-region.

  • size_distribution (dict) – {‘sizes’: […], ‘probs’: […]}. Sampled once per eligible grain to pick that grain’s sub-region count.

  • min_voxels_to_split (int) – Grains with fewer voxels than this are not split.

  • connectivity (int, optional) – 6, 18, or 26. Default 6.

  • random_seed (int, optional) – Seeds a local np.random.Generator: same input always reproduces the same split.

Returns:

  • sub_lgi (np.ndarray) – Same shape as lgi, same nonzero mask. Fresh unique sub-region ids 1..M (ids are NOT related to the original grain ids).

  • sub_to_grain (dict[int, int]) – sub_region_id -> original grain_id it was split from.

  • clusters_dict (dict[int, list[int]]) – original grain_id -> sorted list of sub_region_ids split from it.

upxo.pxtalops.grain_splitting_3d.size_balance_metrics(sizes: Sequence[float]) → Dict[str, float][source]

How far a set of sub-region (or, more generally, any grouped-member) sizes is from a perfectly equal split of their shared total.

Three standard, complementary indicators, each 0 (or 1 for min_max_ratio) at a perfectly even split and moving away from that as the split becomes more unequal:

cv – coefficient of variation (population std / mean) of

the sizes. Since the sizes are assumed to partition a fixed total, their mean already equals the “ideal equal share” (total / n) – so this is exactly the normalised deviation from that ideal share, not merely from whatever the sizes happen to average to. 0 = perfectly equal.

min_max_ratio – smallest / largest size. 1.0 = perfectly equal (every

size identical); lower means the worst-case pair is further apart. Only looks at the two extremes, so it can miss unevenness among the middle sizes.

gini – Gini coefficient of the size distribution (0 =

perfectly equal, larger = more concentrated in a few large sizes). Accounts for the whole distribution rather than just the extremes – complements min_max_ratio.

Parameters:

sizes (sequence of float) – Sizes (e.g. voxel counts) of the sub-regions/members within one parent (grain, PAG, or any other group). A single size (no split occurred, or only one member) returns every indicator at its “perfectly equal” value, since there is nothing to compare.

Return type:

dict with keys ‘cv’, ‘min_max_ratio’, ‘gini’.