upxo.pxtalops.grain_splitting_3d module
Geodesic within-grain splitting for 3D labeled grain structures.
Subdivides each grain in a 3D labeled voxel image (LGI) into one or more connected sub-regions via multi-source BFS (geodesic nearest-seed assignment), seeded from farthest-point-sampled voxels within the grain itself. Standalone and framework-agnostic (pure numpy + stdlib) – takes and returns plain label arrays and dicts, with no dependency on any specific grain-structure class, so it is reusable across UPXO’s various 3D pxtal/grain-structure modules (originally built for, and still used by, the fm_steel_3d PAG/packet pipeline’s “Technique B”, but the algorithm itself is domain-agnostic).
- Guarantees, by construction:
Every output sub-region is a single connected component (a voxel can only be reached by travelling through voxels that belong to the same original grain).
Sub-regions exactly partition their parent grain’s voxels (none lost, none duplicated).
No dependency on inter-grain adjacency/connectivity – splitting one grain never competes with any other grain for resources, unlike clustering-based approaches to forming multi-voxel-region groups.
This is one splitting strategy among several possible ones – geodesic nearest-seed assignment from spread-out (farthest-point-sampled) seeds gives reasonably balanced sub-region sizes (an irregular grain shape or unlucky seed placement can still skew sizes), not an exact equal-volume partition. A different algorithm (e.g. iteratively-rebalanced/weighted seed placement, or a power-diagram-style approach) would be needed for a guaranteed equal-size split; that is a natural direction for future development in this module without changing its calling contract. size_balance_metrics (below) exists precisely to quantify that gap – how far a given split landed from an equal one – so it can be tracked as splitting strategies evolve.
- upxo.pxtalops.grain_splitting_3d.split_grains_geodesic(lgi: numpy.ndarray, size_distribution: Dict[str, Sequence[float]], min_voxels_to_split: int, connectivity: int = 6, random_seed: int | None = None) Tuple[numpy.ndarray, Dict[int, int], Dict[int, List[int]]][source]
Subdivide every grain in
lgiinto 1 or more sub-regions.Grains with fewer than
min_voxels_to_splitvoxels are left whole (single sub-region == the whole grain; this is the correct, intended outcome for a genuinely small grain, not a degenerate/error case). Grains at or above the threshold get a sub-region count sampled fromsize_distribution, then their voxels are partitioned into that many sub-regions via multi-source BFS (geodesic nearest-seed assignment) seeded from farthest-point-sampled voxels within the grain – guaranteed connected sub-regions by construction, since a voxel can only be reached by travelling through voxels that belong to the same original grain.- Parameters:
lgi (np.ndarray) – 3D labeled grain image. 0 = background, 1..N = grain ids. Every nonzero label is assumed to already be a single connected component at the given
connectivity(true for any cc3d.connected_components or Voronoi-argmin output) – this is what guarantees every voxel in a grain ends up assigned to some sub-region.size_distribution (dict) – {‘sizes’: […], ‘probs’: […]}. Sampled once per eligible grain to pick that grain’s sub-region count.
min_voxels_to_split (int) – Grains with fewer voxels than this are not split.
connectivity (int, optional) – 6, 18, or 26. Default 6.
random_seed (int, optional) – Seeds a local np.random.Generator: same input always reproduces the same split.
- Returns:
sub_lgi (np.ndarray) – Same shape as lgi, same nonzero mask. Fresh unique sub-region ids 1..M (ids are NOT related to the original grain ids).
sub_to_grain (dict[int, int]) – sub_region_id -> original grain_id it was split from.
clusters_dict (dict[int, list[int]]) – original grain_id -> sorted list of sub_region_ids split from it.
- upxo.pxtalops.grain_splitting_3d.size_balance_metrics(sizes: Sequence[float]) Dict[str, float][source]
How far a set of sub-region (or, more generally, any grouped-member) sizes is from a perfectly equal split of their shared total.
Three standard, complementary indicators, each 0 (or 1 for min_max_ratio) at a perfectly even split and moving away from that as the split becomes more unequal:
- cv – coefficient of variation (population std / mean) of
the sizes. Since the sizes are assumed to partition a fixed total, their mean already equals the “ideal equal share” (total / n) – so this is exactly the normalised deviation from that ideal share, not merely from whatever the sizes happen to average to. 0 = perfectly equal.
- min_max_ratio – smallest / largest size. 1.0 = perfectly equal (every
size identical); lower means the worst-case pair is further apart. Only looks at the two extremes, so it can miss unevenness among the middle sizes.
- gini – Gini coefficient of the size distribution (0 =
perfectly equal, larger = more concentrated in a few large sizes). Accounts for the whole distribution rather than just the extremes – complements min_max_ratio.
- Parameters:
sizes (sequence of float) – Sizes (e.g. voxel counts) of the sub-regions/members within one parent (grain, PAG, or any other group). A single size (no split occurred, or only one member) returns every indicator at its “perfectly equal” value, since there is nothing to compare.
- Return type:
dict with keys ‘cv’, ‘min_max_ratio’, ‘gini’.