If you see this, something is wrong
First published on Wednesday, Sep 30, 2026 and last modified on Wednesday, Sep 30, 2026 by François Chaplais.
Institute of Material and Process Design, Helmholtz-Zentrum Hereon, Germany and Institute of Production Technology and Systems, Leuphana University Lüneburg, Germany Email
Institute of Material and Process Design, Helmholtz-Zentrum Hereon, Germany
Institute of Material and Process Design, Helmholtz-Zentrum Hereon, Germany
Institute of Material and Process Design, Helmholtz-Zentrum Hereon, Germany
AI for Physical Systems, German Research Center for Artificial Intelligence (DFKI), Germany and Saarland University, Germany
Institute of Material and Process Design, Helmholtz-Zentrum Hereon, Germany and Institute of Production Technology and Systems, Leuphana University Lüneburg, Germany
Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder–decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based–encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer–based decoders on \( \sim80{,}000\) EBSD-derived microstructure dataset to learn a minimal bottleneck, \( z\) . The ViT-FMDiT model (\( z\) =\( 768\) ) reconstructs high-fidelity microstructure images (FID \( 27.86\) , MS-SSIM \( 0.178\) ), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce Meridian, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, failure-aware feasibility prediction, manifold-aware trust regions, and target-aware acquisition. Within a budget of \( 160\) simulations, the ViT-FMDiT and Meridian combination yields the best target-driven objective score, reducing the relative target error by \( 3\) –\( 22\%\) against seven baselines (Dante, TuRBO, BAxUS, CMA-ES, DDOM, SEIKO, DDPO) on the same decoder.
Across science and engineering, the same template recurs: a designer is given a target behaviour and must find a physical object that realises it. A medicinal chemist is given a binding-affinity profile and must find a molecule whose density-functional-theory (DFT) energetics match it [1, 2, 3]; a photonics engineer is given a target far-field radiation pattern and must find a metasurface whose Maxwell-solver response matches it [4, 5]; a crystallographer is given a desired band gap and must find a periodic crystal whose first-principles spectrum matches it [6, 7, 8]. Inverse design asks: given a target property vector \( p^\star\) , find a design \( x^\star \in \mathcal{X}\) such that \( f(x^\star) \approx p^\star\) , where \( f\) is an expensive black-box oracle (e.g., a physical assay or simulator) and a feasibility constraint \( g(x) \le 0\) encodes stability requirements. Three structural obstacles make direct search on \( \mathcal{X}\) intractable. (i) \( \mathcal{X}\) is discrete and non-Euclidean (graphs of atoms, fields of crystallographic orientations), so gradient methods do not apply through a simulator that has no usable adjoint [9]. (ii) most points in \( \mathcal{X}\) are physically invalid, and feasibility cannot be expressed as a simple box, so unguided sampling almost never yields a candidate that oracle will run to convergence [10]. (iii) Each evaluation \( f(x)\) costs hours, with \( 10^2\) –\( 10^3\) evaluation budgets that rules out brute-force sampling, GAN-inversion-by-backpropagation [11, 12], and high-dimensional Bayesian optimization (BO) over the raw space [13, 14, 15]. Prior work side-steps the simulator by training a fast surrogate [16, 17, 6, 7], trading simulator cost for surrogate-fidelity risk that breaks whenever the surrogate is asked to extrapolate past its training distribution.
To ground this framework, we tackle an industrial and scientific challenge that falls entirely out of the ’cheap-oracle’ assumption dominating existing latent-optimization research: target-driven inverse design of polycrystalline Magnesium (Mg) alloy microstructures and textures. The protagonist is the engineer with a target stress–strain curve, eg. a structural designer specifying a Mg-alloy AZ31 component for an automotive crash member [18]; a clinician specifying a bioresorbable Mg–Gd bone screw whose yield strength must match cortical bone over a degradation window [19]; an aerospace designer chasing an anisotropic high-toughness texture in a structural panel [20]. In every case the deliverable is not a microstructure image but an arrangement of grains and crystallographic orientations that, when simulated through a crystal-plasticity (CP) oracle such as Damask [21], returns a property tuple \( (\sigma_y, n, K, \sigma_u)\) within tolerance of the target. Microstructure inverse design is a hard problem because the design space comprises of grains with symmetric-crystallographic orientaions [22, 23] and anisotropic, twin-dominated, and texture-coupled plasticity which causes visually similar microstructures to yield vastly different stress–strain responses [24]; and a single \( 300{\times}300\) CP simulation takes \( \sim\!6\) minutes and risks divergence [21]. Inverse microstructure and texture design are therefore recognised as ill-conditioned, high-dimensional, and data-starved [25, 26].
Why no off-the-shelf recipe applies? To bypass the intractability of raw search space, one could theoretically map the problem into a continuous latent search space with a pretrained generator \( \mathcal{D} : \mathbb{R}^d \to \mathcal{X}\) . In such a setup, decoded samples lie near a learned manifold, the latent prior yields a free box constraint \( \mathcal{Z} = [-B, B]^d\) , and the optimization objective \( J \circ f \circ \mathcal{D}\) is well-defined. This theoretically enables pure target-driven design, where optimal orientations and grain morphologies emerge directly from stress–strain requirements [27]. In practice, however, existing studies assume cheap oracles [28, 1, 29, 30], high-dimensional BO lacks generative priors [13, 14, 31], and recent crystal models [6, 7, 8] are built for sampling, not inverse optimization. None can reliably navigate an expensive, occasionally divergent simulator while enforcing valid material-grain orientation fields.
What is new. The latent representation is an enabling component, not our contribution; the methodological claim concerns the optimizer. Existing active-learning optimizers assume that every query returns a valid value and search generic, isotropic regions. Meridian couples two mechanisms not jointly present in any of the seven baselines: failure-aware feasibility learning, a classifier trained on failed decode/codec/solver evaluations that steers acquisition away from likely divergence, and manifold-aware local search, a surrogate-shaped active-subspace trust region with a latent-shell projection that keeps candidates where the decoder yields valid microstructures. To our knowledge, Co-PiLOT is also the first closed loop coupling an image-domain generative prior to a crystal-plasticity solver under a \( 160\) -simulation budget.
Contributions. We present Co-PiLOT (Constrained Physics-informed Latent Optimization for Target-driven inverse design), a framework that lifts inverse design into a pretrained generative latent space and runs a constrained physics-informed black-box optimizer there.
The 5-tuple abstraction is domain-agnostic: in principle any setting with a pretrained decoder for \( \mathcal{X}\) and a slow oracle scoring \( \mathcal{D}(z)\) is in scope (molecules with a VAE and a DFT oracle [1], photonic devices with an EM solver [34], topology optimization with an FEM oracle [35]), but we demonstrate it only on Mg-alloy microstructures. All code is released (Sec. 6).
High-dimensional latent optimization. Latent-space optimization was popularized by [1], who showed that a learned continuous representation can make structured design amenable to gradient-based and Bayesian optimization. Follow-up latent-BO methods add weighted retraining, trust regions, and constrained or multi-objective variants [28, 29, 30, 36], but they are largely evaluated with cheap neural property predictors rather than expensive physical simulators. High-dimensional BO methods such as TuRBO, BAxUS, and SaasBO improve sample efficiency in large ambient spaces,[13, 14, 31] yet they do not exploit generative priors or explicitly address simulator divergence and structured feasibility, which are central in our setting. Classical black-box optimizers such as CMA-ES remain strong generic baselines [37], although full covariance adaptation scales poorly in high dimensions; we include it in Sec. 5.
Generative inverse design of materials. Generative models have become a standard tool for inverse design in materials and related physical systems [10]. Early materials work used GANs to generate microstructures and couple them to regressors or Bayesian optimization [16, 17], while more recent studies perform inverse design of dual-phase steel microstructures with generative models and BO [38], inverse design of spinodoid architected materials in the small-data regime via BO [39], or use diffusion models for nonlinear mechanical metamaterials and multi-material structures [40, 41]. For polycrystals, even representing microstructure for design remains open [42], and inverse microstructure and texture design rely on informative low-dimensional priors or conditional normalizing flows in active-learning loops [25, 26]. In crystals, CDVAE, DiffCSP, and MatterGen show that modern generative models can capture complex structure distributions [6, 8, 7], but these methods are primarily built for sampling or one-shot conditional generation rather than iterative optimization under a strict expensive-oracle budget.
Black-box optimization around generators. Our work is closest in spirit to methods that place an outer optimization loop around a pretrained generator, as in latent BO for molecules [1, 28], photonic inverse design with neural generators and EM solvers [4, 5, 34], and GAN-based topology optimization with FEM [35, 40]. Diffusion black-box optimizers (DDOM [43]) and reward fine-tuning of diffusion models (DDPO [44], SEIKO [45]) assume offline data or far larger query budgets; optimizing over learned models can exploit adversarial modes [46], and coupling generative models to PDE physics remains frontier work [47]. The key gap is that these settings typically rely on cheaper solvers, smoother objectives, or direct conditional generation. Co-PiLOT instead targets the harsher regime of constrained latent optimization with a slow, occasionally divergent physics solver, and therefore emphasizes calibrated surrogates, feasibility modeling, and robust search in pretrained latent spaces.
The central act of Co-PiLOT is the tight coupling of a pretrained microstructure decoder \( \mathcal{D}\) (Sec. 3.1) with a physics-informed optimizer \( \mathcal{O}\) (Sec. 3.3) via an orientation codec ensuring crystallographic validity (Sec. 3.2). Operating hand-in-hand inside a single loop (Fig. 1), the decoder turns a latent vector \( z{\in}\mathcal{Z}\) into a candidate design, while the optimizer proposes the next \( z\) to query based on a scalar fitness \( J\) returned by an oracle simulator \( f\) (Sec. 3.4).
Problem formulation. Let \( \mathcal{X}\) denote a structured design space (microstructures, molecules, \dots), \( f : \mathcal{X} \to \mathbb{R}^k\) an expensive oracle returning a property tuple \( p = f(x) \in \mathbb{R}^k\) , \( p^\star \in \mathbb{R}^k\) a target vector, \( J : \mathbb{R}^k \to \mathbb{R}\) a scalar objective combining target distance and physics priors, and \( g : \mathcal{X} \to \mathbb{R}^m\) feasibility constraints. Given a pretrained encoder–decoder pair \( (\mathcal{E}, \mathcal{D})\) with \( \mathcal{D} : \mathbb{R}^d \to \mathcal{X}\) , Co-PiLOT maps \( \mathcal{X}\) into the latent box \( \mathcal{Z} = [-B, B]^d\) inherited from the latent prior and solves
(1)
This reformulation has three structural advantages. (i) Deliberate validity as \( \mathcal{D}(z)\) is always a plausible design. (ii) Box constraint gives a natural prior \( \mathcal{Z}\) . (iii) Regularised search space as the decoder restricts the optimizer to a learned manifold, resulting in structured modeling target than direct search over the raw \( \mathcal{X}\) . The materials instantiation (Fig. 1b) dictates the anatomy of this section.
Where physics enters Co-PiLOT. “Physics-informed” follows the physics–ML taxonomies of [48, 68], not PINN-style residual losses. Physics enters at two levels (Fig. 1b). (1) Physics-informed objective: \( J\) is computed from a crystal-plasticity simulation at every iteration (\( J_{V2}\) measures distance to the target properties, \( J_{V1}\) adds a ductility floor, both penalise unphysical grain counts), and Meridian’s feasibility model learns simulator divergence (Secs. 3.3, 3.4); this is knowledge integration via the learning objective [48]. (2) Physics-augmented learned components: the encoder–decoder is trained on codec-encoded EBSD maps whose every pixel is a valid HCP orientation, so the latent manifold is a manifold of physically realisable microstructures, and Meridian’s surrogate is warm-started on \( 100\) crystal-plasticity simulations rather than a cold random prior (Secs. 3.2, 5).
Co-PiLOT is built around a continually pre-trained stack: a ViT-H/14 encoder [49, 50] and diffusion decoders- SDXL [51], SD3.5 [52], and DiT-XL/2 [53], that are pre-trained on massive natural-image datasets. We posit this prior is very effective for microstructure reconstruction on a very thin bottlneck \( z\) . Even a \( d{=}512\) bottleneck (\( {\sim}1500{\times}\) compression) recovers grain morphology, crystallographic colour, and property statistics that are compareable to a from-scratch decoder (Sec. 4).
A single shared encoder architecture. A ViT-H/14 trunk maps \( 512^2{\times}3\) orientation-as-RGB microstructure images to \( 1369\) patch tokens plus [CLS] token. 2B weights are reused with the lower \( 16\) blocks frozen and the upper \( 16\) unfrozen (Fig. 2(a)), retaining the natural-image prior while letting the top of the trunk specialise to EBSD imagery. The [CLS]-only readout is replaced by a four-query Attention Pooler (cross-attention over all \( 1370\) tokens, query self-attention, per-query FFN, then merged) followed by a two-layer MLP that maps to bottleneck \( z\in\mathbb{R}^{d}\) , with \( d{=}512\) as the headline; \( d\in{768,1024}\) are reported as the bottleneck-capacity ablation in Sec. 4.
Three diffusion decoders are each conditioned solely on the bottleneck, \( z\) . The pre-trained backbones are frozen and lightweight adaptors are introduced that translate \( z\) into the native conditioning format of each model (Fig. 2b). FM-DiT (frozen SD3.5 MMDiT [52]): We introduce a LatentToTokens adaptor using \( 16\) queries and \( 5\) QK-normed AdaLN-Zero blocks [53]. This generates \( 16{\times}4096\) tokens for the joint attention layers, while a Pooled Projection head provides the \( 2048\) -D global embedding. SDXL (frozen denoising UNet [51]): The latent \( z\) is mapped to \( 77{\times}2048\) tokens via \( 3\) AdaLN-Zero blocks. We initialize the outer layers with \( 0.10\) Xavier gain, as standard zero-initialization collapses the conditioning heads. A Pooled Projection head maps \( z\) to the \( 1280\) -D embedding, and we apply classifier-free guidance [54]. DiT (DiT-XL/2 [53]): Unlike the others, this model is fully trainable, as freezing the backbone fails during the \( 256{\to}512\) resolution transfer. We repurpose the DiT class-embedding slot to inject \( z\) using a ClassProj head. Additionally, \( 16{\times}1152\) spatial tokens are injected via zero-initialized residuals at blocks \( {0,7,14,21}\) . Architectures for SDXL and DiT are deferred to App. Fig. 8. We also include a pixel-space VQGAN [55] with a separate codebook as a non-\( z\) -bottleneck baseline.
Joint training. Encoder, conditioning adaptors, and backbone (for ViT-DiT) are trained jointly under a composite of latent-space generative terms, a latent-alignment regulariser, and a pixel-space frequency term that is gated to small denoising times \( t\) . The loss function per batch, with \( \hat x_0\) as the clean-latent estimate, \( z\) as the bottleneck, and \( z'\) as its augmentation-twin embedding is given as,
(2)
where \( \mathcal{L}_{\mathrm{diff}}\) is the backbone’s native objective — the rectified-flow velocity loss \( \lVert v_\theta(x_t,t,z) - (\epsilon - x_0)\rVert_2^2\) for ViT-FMDiT [56] and the standard MSE loss for ViT-SDXL and ViT-DiT; \( \mathcal{L}_{\mathrm{rec}} = \mathbb{E}_t\bigl[(1{-}t)^2\,\lVert\hat x_0 - x_0\rVert_2^2\bigr]\) is a timestep-weighted clean-latent reconstruction term that emphasises low-noise steps; \( \mathcal{L}_{\mathrm{con}}\) is an InfoNCE-style contrastive loss over \( {z, z'}\) pairs; \( \mathcal{L}_{\mathrm{vic}}\) is a VICReg variance–covariance regulariser [57] that keeps the bottleneck non-collapsed; \( \mathcal{L}_{\mathrm{freq}} = \lVert \log(1{+}|\mathcal{F}(\hat I)|) - \log(1{+}|\mathcal{F}(I)|)\rVert_1\) is an FFT-amplitude pixel loss with a radial high-frequency mask, evaluated on a small subsample of the batch. The five \( \lambda\) s follow one fixed LossWeightScheduler schedule, reused for every ViT-FMDiT model and never tuned per experiment: \( \lambda_{\mathrm{diff}}\) rises \( 0.5{\to}1.0\) while \( \lambda_{\mathrm{rec}}\) and \( \lambda_{\mathrm{con}}\) decay, and the two regularisers enter with a delayed onset (App. C); perturbing it moves held-out PSNR by at most \( 0.09\) dB (Tab. 6). A two-phase schedule freezes the encoder for the first seven epochs and unfreezes its last \( 16\) blocks at epoch \( 8\) , in step units to absorb the topology change when the optimizer is rebuilt and the cosine schedule recomputed. More details on the implementations can be found in App. C.
Why a dedicated codec? A generic pretrained image decoder emits an RGB tensor; a crystal-plasticity solver expects an HCP orientation field. The codec is the enabling bridge \( \Psi\) in Fig. 1(b): an invertible map between HCP orientation fields and 8-bit RGB images that unlocks the decoder–optimizer loop in Sec. 3.4 to use any image-domain generator as a microstructure prior. Naively mapping Bunge–Euler orientations \( (\varphi_1, \Phi, \varphi_2)\) to RGB triples corrupts the any image-domain decoder data in three ways: periodic wrap-around at \( 0/2\pi\) becomes ringing; For Hexagonically Closed Pack (HCP) fundamental zone, per-grain folding maps physically adjacent grains to opposite quaternions, becoming high-frequency color edges which blurs the decoder output; and recovering the \( q_w\) (quaternions) produces NaNs whenever decoder noise pushes negative square-root. The codec resolves all three through a 5-step encode and a closed vectorised decode (Alg. 1).
Five-step pipeline. Each per-grain orientation is converted to a unit quaternion \( q\in S^3\) [58] under HCP point-group \( D_6\) (\( |\mathcal{S}_{\mathrm{HCP}}|=12\) ). The encode then proceeds: (1) Continuous unfolding: breadth-first-search (BFS) over the grain-adjacency graph replaces each grain’s quaternion by the symmetry equivalent closest to its already-visited parent (Eq. 5), so RGB distance between adjacent grains is proportional to their true misorientation, no artificial fundamental-zone jumps. (2) Class-level anchor: a single per-class anchor \( \bar q_c\) is precomputed via the eigenvalue method of [69] from \( M{=}50\) file-level means; using a class-level anchor enables decoder generated images without per-sample metadata. (3) Frame centring. \( q_g\leftarrow \bar q_c^{-1}\cdot q_g\) concentrates the distribution near identity, away from the \( q_w{=}0\) equator where the double-cover sign flip would reintroduce discontinuities. (4) Stereographic projection. \( S = q_{xyz}/(1+q_w) \in [-1,1]^3\) has the closed-form rational inverse in Eq. 8 (no square roots, numerically stable on all of \( \mathbb{R}^3\) ). (5) Quantisation: \( S\) is linearly mapped to \( [0, 2^b{-}1]^3\) . The per-channel step \( \Delta = 2/(2^b{-}1)\) bounds the worst-case angular error at \( {\approx}0.78^\circ\) . The 8-bit PNG variant is used inside Co-PiLOT, it plugs directly into a pretrained \( \mathcal{D}\) , and its measured \( {\sim}0.6^\circ\) mean error (App. D, Tab. 8) is far below the grain-boundary threshold [32]. Self-segmenting decode is done at inference where no ground-truth label map is available for a generated image. The decoder recovers grains by linking \( 4\) -connected pixels whose max-channel intensity differs by at most \( \tau\) (\( \tau{=}1\) for \( 8\) -bit, \( \tau{=}50\) for \( 16\) -bit image) via a single sparse-graph connected-components pass [33].
With the codec in place, the materials oracle composes into
where the mechanical property tuple \( p = (\sigma_y,\,n,\,K,\,\sigma_u,\,\varepsilon_u)\) collects the \( 0.2\%\) -offset yield stress \( \sigma_y\) , the Hollomon hardening exponent \( n\) , strength coefficient \( K\) , the ultimate tensile stress \( \sigma_u\) , and the uniform strain \( \varepsilon_u\) at \( \sigma_u\) . The grain table emitted by the codec (Sec. 3.2) — per-grain median Euler angles, voxel counts, phase IDs — is written as a Dream.3D representative volume element (RVE) polycrystal that Damask consumes directly. Each candidate is driven through the Damask Grid solver, an FFT-based spectral scheme for periodic RVEs [21], under uniaxial tension along the extrusion direction at a quasi-static strain rate (\( \dot\varepsilon = 10^{-3}\,\mathrm{s}^{-1}\) , \( 25\%\) total nominal strain). The constitutive law is the HCP phenomenological power law of [21] with the Mg-alloy AZ31-calibrated parameter set covering tensile twinning; basal, prismatic, and pyramidal slip deformation mechanisms explained in App. E). Non-converged runs — a small minority, primarily on degenerate decoded grains — are excluded from the surrogate’s training set in all optimizers.
Objective: \( J(p)\) has two complementary forms: a toughness-weighted strength score \( J_{V1}\) that simultaneously rewards high yield \( \sigma_y\) and high plastic work to UTS (proxied by \( \sigma_u\,\varepsilon_u\) ), and a target-driven score \( J_{V2}\) that penalises weighted distance to a user-supplied target \( p^\star=(\sigma_y^\star, n^\star, K^\star, \sigma_u^\star)\) ,
(3)
The \( n_{\min}\) floor in \( J_{V1}\) rejects brittle solutions; both add a saturating grain-count band penalty (App. E) since the constitutive law is grain-size insensitive and the optimizer would otherwise reach the target via a few large favourably-oriented grains or via speckle decodes with thousands of grains. The constants in Eq. (3) (\( \alpha\) , \( \beta\) , \( w_p\) , \( p^\star\) ) are not tuned: they encode the design goal, so changing them changes the question being asked rather than how well it is answered.
We deploy Meridian (Manifold-Embedded Robust Inverse Design via Iterative Acquisition Networks) as the primary optimizer and benchmark it against seven baselines under the same box, seed cache, and budget: Dante [59] (MLP surrogate with tree exploration), TuRBO [13] (GP with a single trust region), BAxUS [14] (GP in a random subspace), CMA-ES [37] and DDOM [43] as optimizer-level comparisons, and SEIKO [45] and DDPO [44], which fine-tune the decoder itself, as framework-level comparisons. Implementations and hyperparameters (Tabs. 10, 11) are in App. F.
Why a dedicated optimizer? Three properties specific to Co-PiLOT shapes the design. The oracle is expensive (\( \sim 15\) min/sim) and a non-trivial fraction of evaluations fail at decode, codec, or solver gates, so the optimizer must model both where the response is uncertain and where it returns a finite value. Pre-trained image-domain decoders place their training mass on a thin spherical shell of the latent box (e.g. \( \| z\| {=} 22.07{\pm}0.05\) for ViT-FMDiT-\( 512\) ), so any perturbations in \( \mathbb{R}^{512}\) leave this shell and waste simulator calls. Dante’s point-estimate MLP and visit-count tree miss all three; the GP-trust-region baselines satisfy uncertainty but neither manifold geometry nor batch diversity. Meridian retains the iterative “surrogate \( \to\) candidate cloud \( \to\) diverse batch” skeleton common to other three optimizers and replaces every component on which it breaks for this problem (Tab. 10).
One round of Meridian. Each outer iteration begins by refitting the surrogate: a shared MLP feature map \( \phi_\theta : \mathbb{R}^{512} \to \mathbb{R}^{16}\) is trained jointly with a logistic feasibility classifier \( g_\psi\) (on all observations, so the failure signal is preserved), after which an exact ARD-Mat
specialChar{39}ern-\( 5/2\) GP is fit on the feasible subset on top of \( \phi_\theta\) to supply calibrated \( (\mu, \sigma^2)\) . From the surrogate we read off an importance-weighted axis vector \( w \in \mathbb{R}^{512}\) via the active subspace of \( \hat C = \mathbb{E}[\nabla\hat\mu\,\nabla\hat\mu^\top]\) (PCA of the class-conditional pool \( \mathcal P\) during cold start) [60]. We then anchor the round on the shell-projected centroid of the top-\( K\) feasible incumbents and draw a Sobol cloud around it whose perturbations are scaled by the trust-region edge \( L_t\) and the importance weights \( w\) on a sparse axis mask, so that proposals respect the manifold’s anisotropy without ignoring rare-direction pockets, and then project the cloud onto the adaptive shell \( \| z\| \in [\mu_{\| X\|} \pm k_\sigma\, \sigma_{\| X\|}]\) , which is measured from the feasible data each round. Acquisition combines a feasibility-gated qLogNoisyExpectedImprovement, \( \alpha_{\mathrm{GP}}(z) = \mathrm{qLogNEI}(z; \mu, \sigma)\, g_\psi(z)\) , a Monte-Carlo EI term, for the target-driven \( V2\) objective, computed through heteroscedastic property heads on \( \phi_\theta\) . The top-\( 256\) candidates ranked by \( \alpha\) are passed to a greedy quality-weighted DPP [61] with a hybrid \( (\phi, z)\) kernel that returns a diverse batch of proposals; the batch is then evaluated through \( \mathcal{D} \to \) Damask , \( L_t\) updated under success/failure cadence, and a restart fired whenever \( L_t\) collapses or best plateau runs overflow \( K_{\mathrm{plat}}\) rounds. Meridian, in this configuration with hyperparameters and more explanation are in App. C.
The encoder–decoder models described in Sec. 3.1 are trained once and then frozen for all downstream optimization experiments in Sec. 5. Expanding on the data corpus of [24], we assemble \( 81{,}756\) training and \( 9{,}084\) held-out test RGB microstructures; this expansion was achieved through geometry and physics-aware oversampling, the details of which are outside the scope of this paper and will be documented in a future publication. All the models from Tab. 1 are trained on \( 4\times\) H200 GPUs, and a standard \( 30\) -epoch run on the training split takes roughly \( 12\) hours. The VQGAN reference, FM-DiT at \( z\in{512,768,1024}\) , and SDXL all use epoch-\( 30\) checkpoints; DiT is substantially heavier because its denoiser backbone is unfrozen, carrying about \( 1.17\) B trainable parameters versus roughly \( 381\) M–\( 391\) M for the FMDiT family, and the comparison therefore uses the latest available epoch-\( 20\) checkpoint from the same training regime. We score reconstruction quality with three computer-vision metrics and two material science metrics. FID is the distribution-level realism score, MS-SSIM captures multi-scale structural similarity, and LPIPS captures perceptual distance [62, 63, 64]. For materials fidelity, we report \( G_{\mathrm{m}}\) for grain recovery (fraction of original grains recovered at IoU \( \ge 0.3\) , capturing grain size and aspect-ratio agreement) and \( \Delta_{\mathrm{ori}}\) for crystallography (mean disorientation between IoU-matched per grains per-class mean quaternions).
| Decoder | \( z\) | FID \( \downarrow\) | MS-SSIM \( \uparrow\) | LPIPS \( \downarrow\) | \( G_{\mathrm{m}}\) \( \uparrow\) | \( \Delta_{\mathrm{ori}}\) \( \downarrow\) |
| No bottleneck for VQGAN and Codec | ||||||
| VQGAN | – | 35.63 | \( 0.876 \pm 0.12\) | \( 0.132 \pm 0.10\) | \( 0.747 \pm 0.16\) | \( 27.619 \pm 10.01\) |
| Codec | – | 0.95 | \( 0.449 \pm 0.28\) | \( 0.368 \pm 0.15\) | \( 0.926 \pm 0.08\) | \( 7.200 \pm 4.26\) |
| SDXL | 512 | 108.28 | \( 0.077 \pm 0.05\) | \( 0.650 \pm 0.08\) | \( 0.370 \pm 0.15\) | \( 56.467 \pm 8.41\) |
| DiT | 512 | \( \mathbf{23.19}^{*}\) | \( 0.147 \pm 0.18\) | \( 0.603 \pm 0.08\) | \( 0.486 \pm 0.12\) | \( 55.363 \pm 10.35\) |
| FM-DiT | 512 | 27.86 | \( 0.169 \pm 0.20\) | \( 0.578 \pm 0.07\) | \( \mathbf{0.501 \pm 0.11}^{*}\) | \( \mathbf{53.921 \pm 11.68}^{*}\) |
| FM-DiT | 768 | 27.86 | \( \mathbf{0.178 \pm 0.21}^{*}\) | \( \mathbf{0.572 \pm 0.07}^{*}\) | \( 0.474 \pm 0.12\) | \( 54.318 \pm 11.16\) |
| FM-DiT | 1024 | 27.08 | \( 0.174 \pm 0.21\) | \( 0.574 \pm 0.07\) | \( 0.475 \pm 0.12\) | \( 54.353 \pm 10.81\) |
ViT-DiT achieves the best FID, while the ViT-FMDiT family (specifically at \( d{=}768\) ) achieves the best \( G_{\mathrm{m}}\) —indicating the strongest grain-level recovery in size and shape—alongside the best MS-SSIM, LPIPS, and \( \Delta_{\mathrm{ori}}\) among bottlenecked decoders (Tab. 1). ViT-FMDiT’s performance is stable across bottleneck size, allowing the optimization loop (Sec. 5) to use a tighter bottleneck without sacrificing fidelity. Conversely, ViT-SDXL struggles to adapt its text-conditioned UNet to a 512-D continuous bottleneck, yielding the worst metrics (FID \( \sim\!108\) ) and visibly washed-out grains in Fig. 3. An unbottlenecked ViT-VQGAN provides an upper bound for reconstruction, while the orientation codec establishes the error floor. The codec’s \( \Delta_{\mathrm{ori}}\) error is safely within CP tolerances. While the bottlenecked decoders preserve macroscopic morphology and class statistics (e.g., AZ31, Mg-10Gd), their orientation space is non-isometric and symmetry-blind. This introduces a \( {\sim}53^\circ\) per-pixel disorientation into the end-to-end round-trip. This residual error acts as a class-preserving reshuffle rather than a physical violation as the decoded grains remain in the HCP fundamental zone, and their texture statistics (quaternion mean, pole-figure spread) matches the target class. Therefore, we carry ViT{-}DiT and ViT{-}FMDiT with \( d\in{512,768,1024}\) into the Co-PiLOT loop as robust generators of statistically valid microstructures rather than exact reconstructors.
E2 tests the central claim : Can Co-PiLOT find statistically valid AZ31 microstructures whose simulated stress–strain summaries match a user-specified mechanical target? We use the optimizer suite from Sec. 3.4 and the trained decoders from Sec. 4. The target, \( p^\star=(\sigma_y^\star,n^\star,K^\star,\sigma_u^\star)=(160\,\mathrm{MPa},0.302,604\,\mathrm{MPa},316\,\mathrm{MPa})\) is set. Each cell uses \( 5\) independent runs and the same \( 100\) (\( z_\mathbf{init}\) <-> Damask evaluation) seed cache for warm starting the surrogate model. The optimizer evaluates \( 40\) iterations with batch size \( 4\) , i.e. \( 160\) new simulator evaluations per run. Average runtime per evaluation on a \( 700\) W H200 is \( 8.46\) hours with a GPU energy estimate of approx \( 5.93\) kWh across the full sweep. Table 2 identifies ViT{-}FMDiT-\( 768\) with Meridian as the strongest target-driven configuration. After 144 evaluations (Fig. 4), it yields the closest overall target match (\( J_{V2}=-0.132\pm0.004\) , \( \sigma_y=159.1\) MPa, \( n=0.299\) , \( \sigma_u=314.0\) MPa), successfully falling within the AZ31 material class. This success highlights a key principle: the optimizer and decoder must align. The \( 768\) -D bottleneck offers the ideal trade-off—more expressive than \( 512\) -D, yet dense enough for local surrogate modeling. Conversely, the \( 1024\) -D bottleneck adds expressivity but becomes too sparse under a \( 160\) -evaluation budget, so larger latents do not automatically improve inverse design. This reinforces our core conclusion: successful inverse design requires an optimizer whose inductive biases match the encoder-decoder’s latent space (see App. A for \( J_{V1}\) ).
| Decoder | Optimizer | \( J_{V2}\uparrow\) | \( \sigma_y\) | \( n\) | \( \sigma_u\) | # sims |
| ViT-DiT | Meridian | \( -0.156 \pm 0.005^*\) | \( 155.9\) | \( 0.305\) | \( 313.9\) | \( 160\) |
| Dante | \( -0.163 \pm 0.009\) | \( 155.8\) | \( 0.305\) | \( 312.9\) | \( 104\) | |
| TuRBO | \( -0.168 \pm 0.007\) | \( 154.1\) | \( 0.301\) | \( 307.7\) | \( 152\) | |
| BAxUS | \( -0.168 \pm 0.009\) | \( 154.8\) | \( 0.303\) | \( 310.6\) | \( 152\) | |
| ViT-FMDiT | Meridian | \( -0.152 \pm 0.004^*\) | \( 156.1\) | \( 0.302\) | \( 312.1\) | \( 96\) |
| Dante | \( -0.161 \pm 0.016\) | \( 157.1\) | \( 0.303\) | \( 314.4\) | \( 156\) | |
| TuRBO | \( -0.167 \pm 0.009\) | \( 154.4\) | \( 0.301\) | \( 308.4\) | \( 160\) | |
| BAxUS | \( -0.162 \pm 0.005\) | \( 155.1\) | \( 0.303\) | \( 311.1\) | \( 144\) | |
| ViT-FMDiT 768 | Meridian\( ^\dagger\) | \( \mathbf{-0.132 \pm 0.004}^*\) | \( 159.1\) | \( 0.299\) | \( 314.0\) | \( 144\) |
| Dante | \( -0.140 \pm 0.004\) | \( 157.7\) | \( 0.301\) | \( 313.1\) | \( 140\) | |
| TuRBO | \( -0.151 \pm 0.009\) | \( 158.5\) | \( 0.306\) | \( 318.5\) | \( 148\) | |
| BAxUS | \( -0.146 \pm 0.004\) | \( 157.7\) | \( 0.305\) | \( 316.5\) | \( 156\) | |
| CMA-ES | \( -0.144 \pm 0.002\) | \( 156.6\) | \( 0.302\) | \( 313.1\) | \( 147\) | |
| DDOM | \( -0.136 \pm 0.008\) | \( 157.9\) | \( 0.301\) | \( 313.1\) | \( 123\) | |
| SEIKO | \( -0.165 \pm 0.009\) | \( 153.5\) | \( 0.304\) | \( 309.5\) | \( 133\) | |
| DDPO | \( -0.170 \pm 0.003\) | \( 152.9\) | \( 0.305\) | \( 309.7\) | \( 144\) | |
| ViT-FMDiT 1024 | Meridian | \( -0.150 \pm 0.005^*\) | \( 157.0\) | \( 0.304\) | \( 314.0\) | \( 152\) |
| Dante | \( -0.168 \pm 0.004\) | \( 153.6\) | \( 0.304\) | \( 308.8\) | \( 124\) | |
| TuRBO | \( -0.156 \pm 0.009\) | \( 157.1\) | \( 0.305\) | \( 315.0\) | \( 140\) | |
| BAxUS | \( -0.164 \pm 0.008\) | \( 155.3\) | \( 0.304\) | \( 311.6\) | \( 124\) | |
| CMA-ES | \( -0.167 \pm 0.005\) | \( 154.0\) | \( 0.307\) | \( 312.6\) | \( 144\) | |
| DDOM | \( -0.154 \pm 0.009\) | \( 155.4\) | \( 0.305\) | \( 313.0\) | \( 111\) | |
| SEIKO | \( -0.174 \pm 0.011\) | \( 152.6\) | \( 0.307\) | \( 310.4\) | \( 111\) | |
| DDPO | \( -0.176 \pm 0.004\) | \( 152.7\) | \( 0.309\) | \( 312.2\) | \( 155\) | |
Extended baselines. On ViT-FMDiT-\( 768\) /\( 1024\) we also ran CMA-ES, DDOM, SEIKO, and DDPO under the same protocol (Tab. 2, lower rows). Meridian stays best (\( -0.132\) /\( -0.150\) ); DDOM (\( -0.136\) /\( -0.154\) ) is within one standard deviation but never ahead, as it models neither feasibility nor trust-region geometry, and surrogate-free CMA-ES trails (\( -0.144\) /\( -0.167\) ) because it must spend simulator calls to learn the landscape. SEIKO and DDPO, which fine-tune the decoder toward average reward, do not beat the warm-start incumbent within \( 160\) evaluations, \( 160\) –\( 360\times\) fewer queries than DDPO’s reference configurations. The middle of the ranking reorders between decoders while the top two do not, so the result reflects the method rather than one bottleneck width. Across four decoders and two objectives, Meridian is best in all eight studies; batch-rule and sensitivity ablations are in App. B.
Limitations. Co-PiLOT is limited by its decoder as it can only search microstructures that the learned prior can express. The decoder is also a class-conditional generator of statistically valid AZ31 microstructures, not a faithful per-pixel orientation reconstructor; although the codec is sub-degree accurate, the encoder–decoder is still non-isometric in crystallographic rotation distance. The Damask oracle is HCP phenopower-law plasticity, no damage, isothermal loading, and the loading paths studied here, therefore, inverse-designed microstructures require multi-loading and experimental validation. This paper does not show transfer to other polycrystalline materials, property sets, or physical domains, and compares optimizers at a single \( 160\) -evaluation budget; harder inverse-design problems are future work. We are addressing these issues by expanding validation and training \( \mathcal{E}\) –\( \mathcal{D}\) with orientation-specific losses such as HCP-symmetry-aware grain-level quaternion losses and texture-statistics penalties.
Highlights and broader impact. Co-PiLOT turns expensive simulator-driven inverse design into active search over a learned generative design space; to our knowledge, no similar end-to-end pipeline previously existed for this regime. The decoders are evaluated by FID, MS-SSIM, LPIPS, grain recovery, and disorientation against ground truth (Tab. 1) and Meridian against seven optimizers under a matched budget plus batch-rule and control-parameter ablations (Tab. 2, App. B). The codec makes image-domain generators usable inside a crystal-plasticity loop, and Meridian combines calibrated uncertainty, feasibility modeling, manifold-aware proposals, and diverse batches to reduce wasted simulator calls. On AZ31 design, ViT-FMDiT-\( 768\) with Meridian gives the strongest target-driven result and remains effective on the toughness objective. The abstraction may extend to other physical systems with learned priors and slow oracles, which remains to be shown. Code: https://github.com/mahishguru/microstructure-encoder-decoder (encoder–decoder), https://github.com/mahishguru/orientation-codec (orientation codec), and https://github.com/mahishguru/meridian (Meridian, Damask harness).
Acknowledgements We thank the anonymous reviewers and the area chair for comments that improved the evaluation of this work.
Funding. We acknowledge Helmholtz-Zentrum Hereon for providing computational infrastructure and institutional support. Additional support was provided by the Joachim Herz Stiftung through an Add-on Fellowship.
Competing interests. The authors declare no competing interests.
This appendix reports the toughness-weighted objective results that complement the target-driven E2 results in Sec. 5. The protocol is unchanged: the shared seed cache is used only for warm start, all curves and tables exclude that cache, and each run receives \( 160\) optimizer-phase simulator evaluations. Average wall-clock runtime for these \( J_{V1}\) runs was \( 8.48\) hours on an H200 machine, corresponding to an estimated \( 5.94\) kWh of GPU energy per run under the same \( 700\) W TDP accounting used in the main text.
| Decoder | Optimizer | \( J_{V1}\uparrow\) | \( \sigma_y\) | \( n\) | \( \sigma_u\) | # sims |
| ViT-DiT | Meridian | \( 0.930 \pm 0.003^*\) | \( 154.3\) | \( 0.308\) | \( 314.0\) | \( 144\) |
| Dante | \( 0.929 \pm 0.010\) | \( 155.6\) | \( 0.305\) | \( 313.4\) | \( 148\) | |
| TuRBO | \( 0.915 \pm 0.007\) | \( 153.8\) | \( 0.306\) | \( 311.3\) | \( 112\) | |
| BAxUS | \( 0.918 \pm 0.002\) | \( 151.8\) | \( 0.312\) | \( 314.2\) | \( 160\) | |
| ViT-FMDiT | Meridian | \( 0.943 \pm 0.011^*\) | \( 156.8\) | \( 0.306\) | \( 316.1\) | \( 124\) |
| Dante | \( 0.923 \pm 0.008\) | \( 155.2\) | \( 0.305\) | \( 313.0\) | \( 152\) | |
| TuRBO | \( 0.931 \pm 0.003\) | \( 154.0\) | \( 0.309\) | \( 314.2\) | \( 100\) | |
| BAxUS | \( 0.933 \pm 0.006\) | \( 156.3\) | \( 0.301\) | \( 311.7\) | \( 100\) | |
| ViT-FMDiT 768 | Meridian\( ^\dagger\) | \( \mathbf{0.961 \pm 0.008}^*\) | \( 157.9\) | \( 0.305\) | \( 316.8\) | \( 112\) |
| Dante | \( 0.954 \pm 0.014\) | \( 159.8\) | \( 0.300\) | \( 316.6\) | \( 140\) | |
| TuRBO | \( 0.948 \pm 0.007\) | \( 156.1\) | \( 0.309\) | \( 317.9\) | \( 160\) | |
| BAxUS | \( 0.953 \pm 0.004\) | \( 157.2\) | \( 0.304\) | \( 315.1\) | \( 116\) | |
| ViT-FMDiT 1024 | Meridian | \( 0.952 \pm 0.005^*\) | \( 156.2\) | \( 0.308\) | \( 317.6\) | \( 160\) |
| Dante | \( 0.923 \pm 0.006\) | \( 153.5\) | \( 0.309\) | \( 314.0\) | \( 156\) | |
| TuRBO | \( 0.944 \pm 0.010\) | \( 157.2\) | \( 0.305\) | \( 315.3\) | \( 140\) | |
| BAxUS | \( 0.936 \pm 0.011\) | \( 155.8\) | \( 0.308\) | \( 316.7\) | \( 128\) | |
| Green = our method (Meridian). \( ^*\) = best per decoder block. \( ^\dagger\) = best overall. | ||||||
The objective changes the design question from hitting a prescribed stress–strain tuple to finding high-strength, high-plastic-work microstructures without dropping below the hardening floor. Figure 6 and Table 3 show that the same decoder–optimizer pairing remains the most effective overall: ViT{-}FMDiT-\( 768\) with Meridian reaches \( J_{V1}=0.961\pm0.008\) , with \( \sigma_y=157.9\) MPa, \( n=0.305\) , \( \sigma_u=316.8\) MPa, and \( 112\) optimizer-phase simulations to the best value. Compared with the target-driven case, the best solution now moves slightly toward higher ultimate strength while preserving hardening, as expected for a toughness-weighted score.
The \( 1024\) -D appendix rows mirror the main-text interpretation. This does not indicate a defective decoder; rather, the wider bottleneck exposes isolated texture pockets that are difficult to model globally from \( 160\) calls. Sparse local perturbations and low-dimensional PCA-aligned subspace moves can exploit such pockets without fitting the whole ambient latent at once, which explains why the best \( J_{V1}\) rows for ViT{-}FMDiT-\( 1024\) come from those search biases. Across both objectives, the consistent conclusion is that ViT{-}FMDiT-\( 768\) gives the best operating point for Co-PiLOT: enough latent capacity to express useful AZ31 textures, but not so much that the expensive-oracle optimizer loses sample efficiency.
Ranking consistency. Our conclusions are ranking statements (which optimizer and which decoder are best), so we test whether the rankings survive variations in the control parameters. Across four decoders and two objectives (Tabs. 2, 3), Meridian attains the best mean objective in all eight studies; under a null of uniformly random ranking among the four optimizers run on every decoder, this occurs with probability \( (1/4)^8 \approx 1.5\times10^{-5}\) . Symmetrically, ViT-FMDiT-\( 768\) is the best decoder in all eight optimizer–objective combinations. The ViT-DiT decoder was trained under a materially different loss-weight configuration from the ViT-FMDiT decoders (\( \lambda_{\mathrm{vic}}\) \( 0.01\) vs. \( 0.05\) , \( \lambda_{\mathrm{freq}}\) \( 0.1\) vs. \( 0.3\) , \( \lambda_{\mathrm{con}}\) starting at \( 0.0\) vs. \( 0.3\) , different onset/ramp epochs), so the optimizer ranking is not an artifact of one loss-weight specification.
DPP versus greedy top-\( k\) batch selection. We swap only Meridian’s DPP batch selector for plain top-\( k\) by acquisition score, holding every other component and the seed cache fixed (\( J_{V2}\) , \( 160\) evaluations, matched seeds; Tab. 4). These are separate matched runs, so values differ slightly from Tab. 2. On ViT-FMDiT-\( 768\) both rules reach the same mean quality (\( J_{V2}=-0.138\) ), so DPP does not raise the ceiling, but it reaches the plateau about \( 11\%\) of the budget earlier (\( 143\) vs. \( 160\) evaluations) with lower run-to-run spread (\( \pm0.007\) vs. \( \pm0.010\) ), and it eliminates a batch-collapse failure mode in which one top-\( k\) seed selected four near-duplicate candidates and made no progress after the first iteration. On ViT-FMDiT-\( 1024\) the two rules tie statistically (\( -0.165\pm0.010\) vs. \( -0.162\pm0.006\) , no collapsed seeds). Under an expensive oracle and a thin-shell latent manifold, batch diversity therefore buys sample efficiency and robustness rather than peak performance.
| Decoder | Batch rule | \( J_{V2}\uparrow\) | \( \sigma_y\) (MPa) | \( n\) | \( \sigma_u\) (MPa) | evals\( \to\) best | collapsed |
| DPP (ours) | \( \mathbf{-0.138 \pm 0.007}\) | \( 157.5\) | \( 0.303\) | \( 315.0\) | \( \mathbf{143}\) | \( \mathbf{0/3}\) | |
| top-\( k\) | \( -0.138 \pm 0.010\) | \( 157.6\) | \( 0.303\) | \( 314.7\) | \( 160\) | \( 1/3\) | |
| DPP (ours) | \( -0.165 \pm 0.010\) | \( 153.6\) | \( 0.305\) | \( 309.7\) | \( 153\) | \( 0/4\) | |
| top-\( k\) | \( -0.162 \pm 0.006\) | \( 154.4\) | \( 0.306\) | \( 312.3\) | \( 146\) | \( 0/4\) | |
Sensitivity to Meridian’s control parameters. Meridian has its own settings: trust-region size and shrink rate, the latent-norm shell bound \( k_\sigma\) , and the batch-diversity term. We grouped repeated ViT-FMDiT-\( 768\) Meridian runs under one identical pair of objectives by the configuration they were produced under (Tab. 5). The evaluation budget is essentially unaffected: under \( J_{V2}\) every setting reaches its plateau at \( 104\) –\( 106\) of the \( 160\) evaluations, and no setting stalls, diverges, or fails to converge. The achieved optimum moves by about one seed-level standard deviation (\( 0.011\) on \( J_{V2}\) , against a per-setting seed spread of \( 0.005\) –\( 0.013\) ). The published setting is best on both objectives although it was tuned only on \( J_{V1}\) . The one setting with a systematic effect is the batch-diversity term, a claimed component of the method ablated above.
| Meridian setting | What differs | \( J_{V1}\uparrow\) | evals | \( J_{V2}\uparrow\) | evals |
| published reference (Alg. 2) | — | \( \mathbf{+0.961 \pm 0.009}\) | \( 58\) | \( \mathbf{-0.136 \pm 0.009}\) | \( 106\) |
| earlier trust-region + shell | compact/fixed trust region, DPP pool \( 64\) , shell \( k_\sigma\in{1.0,2.0}\) , faster/slower shrink | \( +0.944 \pm 0.003\) | \( 53\) | \( -0.147 \pm 0.005\) | \( 104\) |
| DPP \( \to\) greedy top-\( k\) | diversity term removed | \( +0.956 \pm 0.012\) | \( 93\) | \( -0.142 \pm 0.013\) | \( 105\) |
| spread (max\( -\) min) | \( \mathbf{0.017}\) | \( \mathbf{0.011}\) | |||
This appendix records the implementation detail behind the encoder–decoder family that the main text only sketches. The goal is not to restate Sec. 3.1, but to make the exact architectural, training, and validation choices reproducible: the shared encoder, the three conditioning interfaces, the composite training objective, and the bottleneck-capacity ablation.
Shared encoder — exact specification. The encoder is OpenCLIP ViT-H/14 (laion2b_s32b_b79k weights) [49, 50]: \( 32\) transformer blocks, hidden width \( 1280\) , patch size \( 14\) , and input resolution \( 512{\times}512\) , which yields \( 37{\times}37{=}1369\) patch tokens after bicubic interpolation of the original \( 16{\times}16\) positional embedding. We keep CLIP normalisation throughout rather than switching to ImageNet statistics. The freezing pattern is deliberately asymmetric: conv1, cls_token, the positional embedding, ln_pre, and the lower \( 16\) transformer blocks remain frozen, while only the upper \( 16\) blocks are trainable, with gradient checkpointing enabled there. This keeps most of the LAION prior intact while allowing the top of the trunk to adapt to EBSD imagery. The CLS readout is replaced by a multi-query attention pooler with \( 4\) learned queries (initialised \( \mathcal{N}(0,0.02)\) ), \( 8\) attention heads, query self-attention, and a per-query FFN, followed by a merge head \( \mathrm{Linear}(4D \to D)\) . The bottleneck MLP is \( \mathrm{Linear}(1280,1280)\to\mathrm{GELU}\to\mathrm{Linear}(1280,d)\to\mathrm{LayerNorm}\) with \( d \in {512,768,1024}\) .
Conditioning adaptors — exact specifications. All three decoders reuse the same bottleneck but expose different native conditioning interfaces, so each adaptor is designed to match the backbone it plugs into rather than to force a uniform abstraction. FM-DiT. LatentToTokens\( (d \to 16{\times}4096)\) uses \( 16\) learned queries, \( 5\) AdaLN-Zero blocks, and QK-normalised cross-attention against the pooled latent. Its outer \( \mathrm{proj\_out}\) is zero-initialised, which is stable here because flow matching learns a residual velocity field. The pooled head is \( \mathrm{Linear}(d,2048)\to\mathrm{SiLU}\to\mathrm{Linear}(2048,2048)\) , again with a zero-initialised last layer. The adaptor carries roughly \( 11\) M trainable parameters and the pooled head another \( 4\) M. The frozen backbone is stabilityai/stable-diffusion-3.5-medium; its VAE uses \( 16\) channels, \( 8{\times}\) downsampling, and scale factor \( 1.5305\) . Timesteps are sampled as \( t = \sigma(\mathcal{N}(0,1)+\log\,\mathrm{shift})\) with \( \mathrm{shift}{=}3.0\) , matching the published SD3.5 logit-normal schedule, and inference uses FlowMatchEulerDiscrete. SDXL. LatentToTokens\( (d \to 77{\times}2048)\) mirrors the original SDXL text sequence length with \( 77\) learned queries passed through \( 3\) AdaLN-Zero blocks. Here zero-initialisation was empirically unstable, so both the outer projection and the pooled head are initialised with Xavier gain \( 0.10\) ; runs with zero init produced dead conditioning heads. The pooled branch is \( \mathrm{Linear}(d,1280)\to\mathrm{SiLU}\to\mathrm{Linear}(1280,1280)\) feeding add_embedding, and a learned NullToken implements classifier-free guidance dropout with \( p_{\mathrm{drop}}{=}0.1\) . The trainable interface is roughly \( 8\) M parameters in the token adaptor, \( 0.8\) M in the pooled head, plus the \( 512\) -parameter null token. The SDXL UNet backbone remains frozen in bf16; training uses DDPM and inference EulerDiscrete. DiT. DiT-XL/2 (hidden size \( 1152\) , \( 28\) blocks, \( 16\) heads, patch size \( 2\) , VAE sd-vae-ft-mse) is fully trainable because freezing failed under the \( 256{\to}512\) transfer. The class-conditioning slot is repurposed through \( \mathrm{ClassProj}(d \to 1152) = \mathrm{Linear}\to\mathrm{GELU}\to\mathrm{Linear}\) with Xavier initialisation. A pair of forward pre-hooks (_CLIPInjector/_ZeroEmbedder) removes the original repeated class-embedding path and replaces it with a single top-of-stack injection of \( \mathrm{ClassProj}(z)\) . In parallel, LatentToSpatialTokens\( (d \to 16{\times}1152)\) emits \( 16\) spatial tokens that a SpatialCrossFuser inserts, with zero-initialised output projection, at blocks \( {0,7,14,21}\) . Training uses DDPM with \( 1000\) steps and linearly spaced \( \beta\) from \( 10^{-4}\) to \( 0.02\) ; inference uses \( 250\) -step DDIM.
Composite loss and training schedule. Training combines the five terms introduced in Eq. (2) rather than relying on the backbone loss alone. The diffusion term \( \mathcal{L}_{\mathrm{diff}}\) is the rectified-flow velocity MSE \( \lVert v_\theta(x_t,t,z) - (\epsilon - x_0)\rVert_2^2\) for ViT-FMDiT and the standard \( \epsilon\) -prediction MSE \( \lVert\epsilon_\theta(x_t,t,z) - \epsilon\rVert_2^2\) for SDXL and ViT-DiT, always evaluated in the native VAE latent of the corresponding backbone. The reconstruction term is \( \mathcal{L}_{\mathrm{rec}} = \mathbb{E}_t[(1{-}t)^2\lVert\hat x_0-x_0\rVert_2^2]\) , which biases training toward cleaner timesteps; for the velocity parameterisation, \( \hat x_0 = x_t - t v_\theta\) . On top of this, we add a contrastive InfoNCE-style term \( \mathcal{L}_{\mathrm{con}}\) on augmentation twins \( (z,z')\) , a VICReg term \( \mathcal{L}_{\mathrm{vic}}\) [57] with cross-rank all-gather ( var_w \( {=}25\) , cov_w\( {=}1\) , target std \( 1\) ), and a frequency-domain loss \( \mathcal{L}_{\mathrm{freq}}\) given by an \( L_1\) distance between radially weighted log-amplitude FFT spectra of the decoded prediction and the target image. The FFT term is capped to at most two image pairs per batch and only evaluated for \( t<0.7\) so that it improves high-frequency fidelity without destabilising high-noise updates.
Loss-weight schedule. The five \( \lambda\) s are the weightings of the loss components in Eq. (2); they are not tuned independently and none were tuned per experiment. The LossWeightScheduler applies one fixed schedule to every ViT-FMDiT model: (i) \( \lambda_{\mathrm{diff}}\) (\( 0.5{\to}1.0\) ), \( \lambda_{\mathrm{rec}}\) (\( 0.5{\to}0.1\) ), and \( \lambda_{\mathrm{con}}\) (\( 0.3{\to}0.05\) ) are linearly interpolated between fixed endpoints; (ii) the two regularisers are held at constant targets \( \lambda_{\mathrm{vic}}{=}0.05\) and \( \lambda_{\mathrm{freq}}{=}0.30\) but switched on with a delayed onset, i.e. exactly zero for the first epochs and then ramped linearly, so they cannot destabilise the encoder before it has settled. Only four numbers are free per weight (start, end, onset, ramp length); every intermediate value is determined. The remaining constants are not tuned: the VICReg sub-weights are the standard published values, and the frequency-loss sample cap is a memory bound. We use a two-phase encoder schedule: epochs \( 1\) –\( 7\) update only the adaptors and bottleneck, then epoch \( 8\) unfreezes the upper \( 16\) ViT blocks, rebuilds the optimizer, and recomputes the cosine schedule in optimizer-step units. Training runs under PyTorch DDP on \( 4{\times}\) H200 NVL with find_unused_parameters=True to tolerate the topology change; a per-rank skip handshake ensures that NaN-guarded steps are either taken or skipped synchronously across ranks.
Loss-weight sensitivity. We retrained ViT-FMDiT-\( 768\) under deliberate \( \lambda\) perturbations with identical data (a fixed-split subset of \( 39{,}997\) micrographs covering \( 14\) alloys) and identical optimization settings, including a repeat of the reference configuration to measure the run-to-run noise floor, and scored each model on \( 4{,}000\) held-out micrographs (Tab. 6). Every variant converges monotonically (the largest epoch-to-epoch rise in validation loss is \( {<}0.003\) ), and the final validation loss varies by only \( 0.6\%\) across the set. Removing the delayed onset, reducing \( \lambda_{\mathrm{con}}\) six-fold, or changing four weights and both onsets at once shifts PSNR by at most \( 0.09\) dB and grain disorientation by at most \( 0.21^\circ\) , within a small multiple of the reference-versus-repeat noise floor; no configuration diverges or collapses. In archived runs spanning a \( 75\) -fold range of \( \lambda_{\mathrm{con}}\) , a \( 10\) -fold range of \( \lambda_{\mathrm{diff}}\) , and \( \lambda_{\mathrm{vic}}\in[0,0.2]\) , every configuration likewise converged monotonically.
| Configuration | \( \lambda\) change | Val. loss \( \downarrow\) | PSNR \( \uparrow\) | SSIM \( \uparrow\) | LPIPS \( \downarrow\) | Disor. (\( ^\circ\) ) \( \downarrow\) |
| reference | committed schedule | \( 0.4620\) | \( 12.47\) | \( 0.303\) | \( 0.592\) | \( 59.15\) |
| reference (repeat) | none (noise floor) | \( 0.4621\) | \( 12.45\) | \( 0.298\) | \( 0.594\) | \( 58.99\) |
| no delayed onset | regularisers from epoch \( 0\) | \( 0.4595\) | \( 12.39\) | \( 0.285\) | \( 0.587\) | \( 58.94\) |
| weak-reg + late onset | \( 4\) weights + both onsets | \( 0.4603\) | \( 12.48\) | \( 0.299\) | \( 0.589\) | \( 59.05\) |
| low \( \lambda_{\mathrm{con}}\) | \( \lambda_{\mathrm{con}}\times 1/6\) | \( 0.4604\) | \( 12.39\) | \( 0.291\) | \( 0.590\) | \( 59.08\) |
| spread (max\( -\) min) | \( \mathbf{0.0026}\) | \( \mathbf{0.09}\) | \( \mathbf{0.018}\) | \( \mathbf{0.007}\) | \( \mathbf{0.21}\) | |
Bottleneck-capacity ablation (\( d\in{512,768,1024}\) ). Bottleneck width is the most controlled ablation in this family because it leaves the decoder, encoder, conditioning path, and loss schedule unchanged. Under that constraint, held-out ViT-FMDiT PSNR improves monotonically with width, from roughly \( 14.0\) dB at epoch \( 2\) for \( d{=}768\) to roughly \( 16.9\) dB at epoch \( 2\) for \( d{=}1024\) . This is consistent with the main-text choice of \( d{=}512\) as a deliberately severe bottleneck for the latent-BO experiments rather than as the reconstruction-optimal point. Final-epoch reconstruction metrics for all three widths appear in Tab. 1.
| Hyperparameter | ViT-FMDiT | ViT-DiT | ViT-SDXL |
| Backbone | SD3.5 MMDiT (\( {\sim}2.5\) B) | DiT-XL/2 (\( {\sim}0.7\) B) | SDXL UNet (\( {\sim}2.6\) B) |
| Backbone state | frozen | trainable | frozen (LoRA \( r{=}16\) ) |
| VAE | SD3.5 (16ch, \( s{=}1.5305\) ) | sd-vae-ft-mse (4ch) | SD-XL VAE |
| Bottleneck \( d\) (head./abl.) | \( 512/{768,1024}\) | \( 512/{768,1024}\) | \( 512\) |
| LatentToTokens (#q / depth) | \( 16 / 5\) | \( 16\) spatial / – | \( 77 / 3\) |
| Conditioning dim | \( 4096\) | \( 1152\) | \( 2048\) |
| Pooled-projection dim | \( 2048\) | — | \( 1280\) |
| Spatial injection blocks | — | \( {0,7,14,21}\) | — |
| Outer-proj init | zero | Xavier (\( 0.10\) ) | Xavier (\( 0.10\) ) |
| Trainable interface | \( {\sim}0.6\%\) | \( 100\%\) | \( {\sim}0.4\%\) |
| Per-GPU/accum/eff. batch | \( 48/3/576\) | \( 144/1/576\) | \( 144/2/1152\) |
| Train scheduler | FM-Euler | DDPM (lin., \( 1000\) ) | DDPM |
| Inference scheduler (\( N\) steps) | FM-Euler (\( 50\) ) | DDIM (\( 250\) ) | EulerDiscrete |
| Timestep distribution | logit-N, shift \( 3\) | uniform | uniform |
| CFG dropout | — | — | \( p_{\mathrm{drop}}{=}0.1\) |
This section expands the codec analysis from Sec. 3.2 and makes explicit the geometric choices that let a generic image decoder serve as a microstructure prior. The thread is simple: every step is chosen to preserve local orientation structure while keeping the RGB image numerically stable for downstream generative models.
Quaternion parameterisation and HCP symmetry. We first convert each per-grain orientation from Bunge–Euler angles (\( ZXZ\) ) to a unit quaternion \( q=(q_x,q_y,q_z,q_w)\in S^3\) [58]. Because quaternions form a double cover (\( q\equiv -q\) ), we fix the sign by enforcing the hemisphere convention \( q_w\ge 0\) . The relevant crystal symmetry is the HCP point group \( D_6\) with \( |\mathcal{S}_{\mathrm{HCP}}|=12\) , generated by the six \( c\) -axis rotations \( R_z(k\pi/3)\) , \( k=0,…,5\) , and the six compositions \( R_z(k\pi/3)\cdot R_x(\pi)\) . In quaternion coordinates \( R_z(\theta)=(\sin(\theta/2)\,\hat z,\cos(\theta/2))\) under the standard Hamilton product. Two orientations are therefore physically identical whenever they differ by a symmetry element \( s_k \in \mathcal{S}_{\mathrm{HCP}}\) .
Step 1 — continuous unfolding (anchored BFS). Let \( \mathcal{N}=(\mathcal{V},\mathcal{E})\) denote the grain adjacency graph from Dream.3D, with \( |\mathcal{V}|=G\) grains and \( \mathcal{E}\) the set of \( 4\) -connected boundary pairs. We choose the BFS root deterministically as the largest grain,
(4)
and align its quaternion to the class anchor (Step 2) by \( q_{g^\star} \leftarrow \arg\max_{s,\sigma}\, \sigma\,\langle s\cdot q_{g^\star},\, \bar{q}_c\rangle\) over \( s\in\mathcal{S}_{\mathrm{HCP}}\) and \( \sigma\in{\pm 1}\) , where \( \langle\cdot,\cdot\rangle\) is the \( \mathbb{R}^4\) inner product. Using the largest grain as the seed shortens the average BFS path length, and aligning that single seed to \( \bar q_c\) keeps the unfolded tree on a common hemisphere so that identical orientations map to consistent colours across the class. Each remaining grain then inherits, from its already visited neighbour \( g_p\) , the symmetry equivalent and sign with maximal inner product:
(5)
Each edge requires \( 24\) dot products, so the encode remains \( \mathcal{O}(|\mathcal{E}|)\) in the boundary count. After this pass, quaternion distance between adjacent grains tracks their physical misorientation directly, without the artificial jumps induced by per-grain folding to a fundamental zone.
Step 2 — class anchor via Markley averaging. For each class \( c\) , the anchor \( \bar q_c\) is the \( L_2\) -optimal mean of \( M{=}50\) representative file-level mean quaternions \( {m_j}_{j=1}^M\) . Following [69], it is the leading eigenvector of
(6)
The point of using a class-level rather than per-sample anchor is that generated images arrive without per-sample metadata. The cost is a mild long-range colour drift induced by within-class spread, but at \( M{=}50\) that drift is still comfortably inside the capacity of the encoder bottleneck.
Steps 3–4 — frame centring and stereographic projection. Once the anchor is factored out via \( q_g \leftarrow \bar q_c^{-1}\cdot q_g\) , the distribution concentrates near identity (\( q_w\approx 1\) , \( \left\lVert q_{xyz}\right\rVert \ll 1\) ). Enforcing \( q_w\ge 0\) , we project to
(7)
which is \( C^\infty\) on the open hemisphere \( q_w>0\) . The inverse map is closed form,
(8)
and uses only rational arithmetic. This is the crucial numerical simplification: there is no square root, so decoder noise cannot drive the inverse into NaNs. As \( \left\lVert S\right\rVert\to\infty\) , \( q_w\) approaches \( -1\) smoothly rather than catastrophically.
Step 5 — quantisation error bound. For a centred unit quaternion near identity, let \( S=q_{xyz}/(1+q_w)\) and perturb it by \( \delta S\) . To first order, \( \delta q_{xyz}\approx (1+q_w)\,\delta S\) while \( \delta q_w\) is second order. The induced angular error is therefore \( \delta\theta = 2\arcsin(\left\lVert\delta q_{xyz}\right\rVert) \approx 2(1+q_w)\left\lVert\delta S\right\rVert \approx 2\left\lVert\delta S\right\rVert\) near identity. With quantisation step \( \Delta = 2/(2^b{-}1)\) , this yields the worst-case bound \( \delta\theta \lesssim 0.78^\circ\) for \( b{=}8\) and \( 0.004^\circ\) for \( b{=}16\) . The measured PNG mean in Tab. 8 is \( 0.6^\circ\) , in line with that estimate.
| Class | Boundary F1 | Grain-size \( r\) | PNG-vs-TIFF mean (\( ^\circ\) ) | PNG-vs-TIFF max (\( ^\circ\) ) |
| ME21_extruded | \( 0.9999\) | \( 0.9997\) | \( 0.61\) | \( 1.35\) |
| Mg{-}10Gd_extruded | \( 1.0000\) | \( 0.9993\) | \( 0.69\) | \( 1.30\) |
| AZ31_extruded_HT | \( 0.9999\) | \( 0.9991\) | \( 0.61\) | \( 1.37\) |
Why these choices. The design logic is now clear. The rational inverse in (8) replaces the unstable recovery \( \sqrt{1-q_x^2-q_y^2-q_z^2}\) , which fails as soon as decoder noise makes the radicand negative. Continuous unfolding preserves the proportionality between RGB distance and true misorientation, keeping the encoded image in the low-frequency regime that pretrained image backbones model well. The class-level Markley anchor is not mathematically ideal, but it is the only option compatible with generated images that lack per-sample metadata. Finally, we use \( b{=}8\) inside Co-PiLOT because pretrained ViT and diffusion backbones expect 8-bit RGB, and the resulting \( \sim 0.6^\circ\) floor remains below both typical EBSD angular resolution and the Read–Shockley low-angle threshold.
This section records the constitutive model and post-processing pipeline actually used by the oracle. The main text only needs the fact that Damask returns a stress–strain response; here we spell out the kinematics, hardening law, and property extraction so the simulation layer is unambiguous.
Kinematics and resolved shear stress. We use the standard multiplicative split \( \mathbf{F}=\mathbf{F}_e\mathbf{F}_p\) , which separates elastic lattice stretch from plastic shear [65]. Plastic flow is written as the sum over all active slip and twin systems, \( \mathbf{L}_p = \sum_\alpha \dot\gamma^\alpha\, \mathbf{s}^\alpha \otimes \mathbf{m}^\alpha\) , with \( \mathbf{s}^\alpha\) and \( \mathbf{m}^\alpha\) the unit slip or twin direction and plane normal. The resolved shear stress on system \( \alpha\) is then \( \tau^\alpha = \boldsymbol{\sigma}:(\mathbf{s}^\alpha\otimes\mathbf{m}^\alpha)\) .
Power-law flow rule. Shear rates follow a phenomenological power law, separately on slip (\( sl\) ) and twin (\( tw\) ) systems:
(9)
with reference rates \( \dot\gamma_0^{sl/tw}\) , rate-sensitivity exponents \( n_{sl/tw}\) , and critical resolved shear stresses \( \xi^\alpha_{sl/tw}\) . Twinning contributes only for \( \tau^\alpha>0\) , reflecting the polarity of twin shear in HCP crystals [66].
Hardening evolution. The CRSS evolves through self- and latent-hardening interactions, \( \dot\xi^\alpha = \sum_\beta h^{\alpha\beta}\,|\dot\gamma^\beta|\) , where \( h^{\alpha\beta}\) is the system-interaction matrix. Slip systems use saturation-type hardening with explicit slip–twin coupling,
(10)
with initial slip-hardening modulus \( h_0^{sl}\) , saturation CRSS \( \xi^\infty_{sl}\) , hardening exponent \( a_{sl}\) , and slip–twin coupling factor \( f^{sl\text{-}tw}_{sat}\) . Twin systems harden symmetrically through twin–twin and twin–slip channels,
(11)
The model is intentionally phenomenological: hardening is encoded through these scalar moduli rather than through dislocation-density or twin-volume-fraction state variables [65]. That choice keeps the oracle tractable while preserving the slip–twin asymmetry that matters most in HCP magnesium [66]. Although damage-augmented variants exist [67], we do not use them here. Instead, every run uses the AZ31-calibrated parameter file distributed with Damask (AZ31_Phenopower.yaml), with basal, prismatic, and pyramidal-\( \langle a\rangle\) /-\( \langle c{+}a\rangle\) slip plus tensile twinning, under isothermal uniaxial tension along the extrusion direction and traction-free transverse faces.
Hollomon fit and property extraction. Post-processing is straightforward but consistent across all runs. We read the HDF5 output, compute volume-averaged Cauchy stress and logarithmic strain along the loading axis at every increment, identify the elastic–plastic transition by the \( 0.2\%\) offset rule, and fit the plastic branch in log–log coordinates with the Hollomon law \( \sigma = K\varepsilon_p^n\) . The yield stress \( \sigma_y\) is the offset intercept, \( \sigma_u\) the maximum stress, and \( (K,n)\) the fitted Hollomon parameters. Runs that do not converge, contain NaN stresses, or produce fewer than \( 50\) converged increments are discarded and logged.
| Parameter | Value |
| Loading mode | uniaxial tension along ED |
| Total time \( t\) (s) | \( 250\) |
| Increments \( N\) | \( 1250\) |
| Strain rate \( \dot{\varepsilon}\) (s\( ^{-1}\) ) | \( 1.0 \times 10^{-3}\) |
| Output frequency \( f_{\mathrm{out}}\) | every increment |
| Constitutive law | phenopowerlaw (HCP) |
| Slip families | basal, prismatic, pyr. \( \langle a\rangle\) , pyr. \( \langle c{+}a\rangle\) |
| Twinning | tensile twin only |
| Phase parameter file | AZ31_Phenopower.yaml |
| Failure threshold | \( {<}50\) converged increments \( \Rightarrow\) drop |
This appendix gathers the implementation detail for the optimizer benchmark that would be distracting in the main text: the pipeline-level adaptations applied uniformly across methods, the concrete baseline settings, the compact design comparison in Tab. 10, and the full hyperparameter table in Tab. 11. All values are those committed in config/az31_*.yaml and copilot/optimizers/ of the released code; no per-decoder, per-objective, or per-seed retuning was introduced after the sweep design was fixed.
Surrogate masking of failed evaluations. The decode\( \to\) codec\( \to\) Damask pipeline can reject a candidate for several reasons: the decoded image may collapse to blank or speckle structure, it may fail the chromatic-spread gate, it may produce fewer than \( g_{\min}{=}40\) grains, or the simulator may fail to converge. Following standard latent-BO practice [13, 14], we log such cases for bookkeeping but exclude them from the regression surrogate in Dante, TuRBO, BAxUS, and Meridian. A constant penalty would pollute posterior variance without adding useful structure and would do so symmetrically across baselines. Meridian is the only method that still retains these failures explicitly through its feasibility classifier \( g_\psi\) .
Trust-region restart cap. The second global intervention is the restart cap. Under the published TuRBO rule \( \tau_{\mathrm{fail}}=\lceil d/q \cdot \alpha \rceil\) , the failure tolerance becomes \( 192\) for TuRBO and Meridian at \( (d,q,\alpha)=(512,4,1.5)\) , and \( 128\) for BAxUS at its initial subspace dimension \( d_0=128\) . In a \( 200\) -evaluation regime that effectively disables restarts and turns all three methods into one-shot local searches. We therefore impose a shared guardrail \( \tau_{\mathrm{fail}}^{\max}=8\) on the three trust-region methods while leaving the published \( \alpha\) unchanged. The point is not to strengthen any one method, but to make the documented restart logic reachable at all. Dante has no trust-region mechanism and is unchanged.
For Dante, we use the authors’ Neural Tree Explorer (NTE): an iterative root-and-cloud DUCB tree search paired with a deep MLP surrogate. The surrogate has hidden widths \( [1024,512,256]\) , dropout \( p=0.1\) , and is trained for \( 200\) epochs with AdamW (lr \( 10^{-3}\) , weight decay \( 10^{-4}\) , batch size \( 32\) , validation split \( 0.1\) , early-stopping patience \( 20\) ). Each round expands \( 200\) leaves with branching factor \( 16\) and depth-decayed Gaussian leaves using \( \sigma_{\mathrm{init}}=0.05\) , \( \sigma_{\mathrm{decay}}=0.995\) , DUCB constant \( c_0=0.1\) , and smoothing \( \rho=0.5\) . Since Dante has no native restart mechanism, we leave its visit-accumulation behaviour untouched.
TuRBO is run as a single trust-region GP method with an ARD-Mat
specialChar{39}ern-\( 5/2\) kernel and Thompson sampling over \( 5{,}000\) Sobol candidates. The success/failure cadence is the standard one, \( \tau_{\mathrm{succ}}{=}3\) and \( \tau_{\mathrm{fail}}{=}\lceil d/q \cdot \alpha\rceil\) with \( \alpha=1.5\) , but with the shared cap \( \tau_{\mathrm{fail}}^{\max}{=}8\) described above. Trust-region bounds are \( L\in[0.005,1.6]\) with \( L_{\mathrm{init}}=0.8\) , and collapse triggers a Sobol cold start in \( \mathcal{Z}\) . We also activate a plateau-restart guardrail (B1): after \( K_{\mathrm{plat}}{=}5\) non-improving batches, the method is restarted even if the trust region has not yet formally collapsed. Without this extra trigger, the published shrink schedule would require about seven halvings to reach \( L_{\min}\) , which does not fit inside our \( 40\) -iteration budget.
BAxUS uses the same acquisition and trust-region cadence as TuRBO, but fits its GP in a low-dimensional embedded subspace. We begin at the authors’ default target dimension \( d_0=10\) and replace the usual random sparse \( \pm 1\) embedding with a PCA-of-seeds basis (B2): the projector columns are the leading PCA directions of the class-conditional latent pool \( \mathcal P\) , truncated to \( k_{\mathrm{eff}}=\min(d_0,\mathrm{rank}(\mathcal P))\) . This keeps lifted proposals aligned with the encoded seed support. We also enable reject feedback to the GP (wants_reject_feedback = True) so that soft-floor codec failures are treated as informative evidence about bad directions rather than discarded.
These four baselines share the decoder, latent box, \( 100\) -simulation seed cache, batch size \( q{=}4\) , \( 160\) -evaluation budget, and five seeds with the methods above, and are run on ViT-FMDiT-\( 768\) and -\( 1024\) . CMA-ES and DDOM are direct optimizer-level comparisons to Meridian; SEIKO and DDPO modify the decoder itself and therefore serve as framework-level comparisons to Co-PiLOT. CMA-ES [37] maintains a Gaussian search distribution over \( \mathcal{Z}\) , evaluates samples, and adapts its mean and covariance toward high-scoring candidates. It has no surrogate of the design space, so everything it learns is paid for in simulator calls, whereas Meridian fits a predictive model to every past simulation before proposing. DDOM [43] learns a generative model over designs conditioned on objective value and samples toward high values; adapted from its original offline setting to our fixed-budget loop, it is retrained each round to lean toward the best result seen so far. It proposes plausible candidates but checks neither whether they will run in the simulator nor localises search around the incumbent, the two capabilities Meridian adds through \( g_\psi\) and the trust region. SEIKO [45] repeatedly fine-tunes the decoder toward high reward while keeping it close to its previous version for stability. It improves the generator’s average tendency to produce good designs rather than exploiting the single best design found so far, which limits it under our budget. DDPO [44] treats denoising as a sequential decision process and fine-tunes the decoder by policy gradient on the reward. It is data-hungry: its reference configurations use \( 25{,}600\) –\( 57{,}600\) reward queries, \( 160\) –\( 360\times\) our budget.
| Design axis | Dante [59] | TuRBO [13] | BAxUS [14] | Meridian (ours) |
| Surrogate | MLP point estimate | ARD-Matérn-\( 5/2\) GP | ARD-Matéern-\( 5/2\) GP in subspace | deep-kernel GP |
| Target-aware signal | scalar-\( Y\) only | scalar-\( Y\) only | scalar-\( Y\) only | property-head EI for \( V2\) |
| Feasibility handling | none | none | reject feedback to GP | shared-trunk classifier \( g_\psi\) |
| Proposal geometry | global tree expansion | isotropic trust region | low-dim subspace trust region | anisotropic trust region |
| Manifold prior | none | none | PCA-of-seeds embedding | active subspace + adaptive shell |
| Batch strategy | top-\( k\) predicted | max-posterior sampling | max-posterior sampling | quality-weighted DPP |
| Restart rule | none | Sobol restart + plateau cap | subspace doubling | class-pool or mini-MCTS restart |
| Final-phase behavior | monotone | monotone | monotone | explicit polish phase |
| Model | MLP \( [1024,512,256]\) |
| Acquisition | DUCB |
| Candidates | \( 200\) leaves |
| Geometry | branching \( 16\) tree |
| Training | AdamW, \( 200\) epochs |
| lr / wd | \( 10^{-3}\) / \( 10^{-4}\) |
| Regularisation | dropout \( 0.1\) |
| Exploration | \( \sigma_{\mathrm{init}}=0.05\) |
| Decay | \( \sigma_{\mathrm{decay}}=0.995\) |
| Score | \( c_0=0.1\) , \( \rho=0.5\) |
| Restart | none |
This subsection records the implementation choices that are intentionally omitted from Sec. 3.4. The emphasis here is not on re-deriving the main algorithm, but on documenting the concrete engineering choices that make the method stable in the expensive-oracle regime. All numerical settings appear in Tab. 11 and the committed values in config/az31_*_meridian_v2.yaml. The surrogate stack is a deep-kernel model in which a two-hidden-layer MLP \( \phi_\theta : \mathbb{R}^{512} \to \mathbb{R}^{16}\) (widths \( [512,128]\) , LayerNorm + GELU) is trained jointly with the feasibility classifier \( g_\psi\) on all observations, preserving the decode/codec/simulator failure signal, and the exact GP is then refit only on the feasible subset. For target-driven \( V2\) , the same feature trunk also feeds a heteroscedastic linear property head \( h_p : \mathbb{R}^{16} \to \mathbb{R}^{2\cdot 5}\) that predicts \( (\mu,\log\sigma^2)\) for \( (\sigma_y,\sigma_u,n,K,n_{\mathrm{grains}})\) in standardised units. The acquisition used in practice is therefore
(12)
where \( \tilde\alpha_{\mathrm{GP}}\) and \( \tilde\alpha_{\mathrm{prop}}\) are \( z\) -standardised scores and \( \tilde\alpha_{\mathrm{prop}}\) is estimated by Monte-Carlo EI with \( 64\) samples. The property-head linear layer is reset every \( K_{\mathrm{reset}}{=}10\) refits to avoid variance collapse on a saturated cache, while the shared feature trunk is warm-started throughout.
Proposal generation is tied explicitly to the latent manifold learned by the decoder. After the first two rounds, search anisotropy is read off from the active-subspace matrix
(13)
whose leading eigenspace is truncated at \( 95\%\) cumulative energy and reprojected to ambient axis weights \( w\in\mathbb{R}^{512}\) ; during cold start, that role is played instead by PCA on the class-conditional latent pool \( \mathcal P\) . Around the shell-projected centroid of the top-\( K{=}5\) feasible incumbents, Meridian draws a Sobol cloud of \( N{=}4096\) candidates, perturbs only a sparse mask of axes with \( \Pr[M_{ij}{=}1]=20/d\) , and replaces a fraction \( \rho_{\mathrm{diff}}{=}0.20\) with isotropic Gaussian moves at scale \( 0.5L_t\bar w\) so that diffuse high-value pockets are not ruled out a priori. Most proposals are then clipped to the adaptive shell
(14)
estimated from the feasible samples, while a small exempt fraction \( \rho_{\mathrm{exempt}}{=}0.15\) is left off-shell in case the support genuinely needs to widen. The trust-region edge evolves in the range \( L_t\in[0.05,1.0]\) with \( L_{\mathrm{init}}{=}0.6\) under the same success/failure cadence used by TuRBO/BAxUS.
Batch construction and restart are likewise tuned for the expensive-oracle regime rather than for formal neatness. From the top-\( 256\) candidates by acquisition value, a greedy DPP selects the final batch with kernel
(15)
so diversity is enforced mainly in feature space without giving up direct latent-space separation. For \( V2\) , one slot is overwritten by a short property-gradient seed and one by a jittered incumbent-polish seed. Past \( T_{\mathrm{pol}}{=}35\) of \( T{=}40\) rounds, the diffuse fraction, shell exemption, and latent-RBF blend are all set to zero and \( L_t\) is capped at \( 0.1\) , yielding a deliberately local polishing phase. A restart is triggered either by trust-region collapse (\( L_t<L_{\min}\) ) or by a plateau of \( K_{\mathrm{plat}}{=}5\) rounds; the next anchor is drawn from the class pool or from a small surrogate-side mini-MCTS, the property heads are reset, and termination occurs after \( r_{\max}{=}3\) unsuccessful restarts. Algorithm \thealgorithm. This narrative description is the full supplementary specification referenced from the main text.
We use each baseline at the configuration recommended by its authors, with the two pipeline-level adaptations described above (failure masking; restart cap) applied uniformly across the four latent-BO optimizers, plus the two manifold-awareness adjustments to BAxUS (PCA-of-seeds embedding, soft-floor reject feedback) detailed in the BAxUS subsection above. The committed hyperparameter values are listed in Tab. 11.
| Checklist item | Where addressed | Asset |
| Code released | Secs. 1, 6 | encoder–decoder: https://github.com/mahishguru/microstructure-encoder-decoder ; orientation codec: https://github.com/mahishguru/orientation-codec ; Meridian and Damask harness: https://github.com/mahishguru/meridian |
| Decoder hyperparameters | App. C | Tabs. 7, 6 |
| Simulator hyperparameters | App. E | Tab. 9 |
| Optimizer hyperparameters | App. F | Tab. 11 |
| Sensitivity / ablations | App. B | Tabs. 4, 5 |
| Compute disclosed | Sec. 4; Sec. 5; App. A | — |
| Limitations stated | Sec. 6 | — |
| Training Dataset | Sec. 4 | https://zenodo.org/records/23036836 |
[1] Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules ACS Central Science 2018 4 2 268–276 10.1021/acscentsci.7b00572
[2] Deep learning workflow for the inverse design of molecules with specific optoelectronic properties Scientific Reports 2023 13 1 20031 10.1038/s41598-023-45385-9
[3] GEOM, energy-annotated molecular conformations for property prediction and molecular generation Scientific Data 2022 9 1 185 10.1038/s41597-022-01288-4
[4] Inverse design in nanophotonics Nature Photonics 2018 12 11 659–670 10.1038/s41566-018-0246-9
[5] Inverse Design of Metasurfaces Based on Coupled-Mode Theory and Adjoint Optimization ACS Photonics 2021 8 8 2265–2273 10.1021/acsphotonics.1c00100
[6] Crystal Diffusion Variational Autoencoder for Periodic Material Generation International Conference on Learning Representations 2022
[7] A generative model for inorganic materials design Nature 2025 639 8055 624–632 10.1038/s41586-025-08628-5
[8] Crystal Structure Prediction by Joint Equivariant Diffusion on Lattices and Fractional Coordinates Workshop on ''Machine Learning for Materials'' ICLR 2023 2023
[9] End-to-end differentiability and tensor processing unit computing to accelerate materials' inverse design npj Computational Materials 2023 9 1 121 10.1038/s41524-023-01080-x
[10] Inverse design with deep generative models: next step in materials discovery National Science Review 2022 9 8 nwac111 08 10.1093/nsr/nwac111
[11] Inverting The Generator Of A Generative Adversarial Network (II) 2018
[12] Towards Understanding the Mechanisms of Classifier-Free Guidance, Xiang Li and Rongrong Wang and Qing Qu, The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026, https://openreviewṅet/forum?id=bRAm7A02Qm
[13] Scalable Global Optimization via Local Bayesian Optimization Advances in Neural Information Processing Systems 2019 H. Wallach and H. Larochelle and A. Beygelzimer and F. d' Alché-Buc and E. Fox and R. Garnett 32 Curran Associates, Inc.
[14] Increasing the Scope as You Learn: Adaptive Bayesian Optimization in Nested Subspaces Advances in Neural Information Processing Systems 2022 Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho
[15] Bayesian Optimization for Materials Design 45–75 Springer International Publishing 2015 Dec 10.1007/978-3-319-23871-5_3
[16] Microstructural Materials Design Via Deep Adversarial Learning Methodology Journal of Mechanical Design 2018 140 11 111416 10 10.1115/1.4041371
[17] Improving direct physical properties prediction of heterogeneous materials from imaging data via convolutional neural network and a morphology-aware generative model Computational Materials Science 2018 150 212-221 https://doi.org/10.1016/j.commatsci.2018.03.074
[18] Mathaudhu, Suveen N. and Luo, Alan A. and Neelameggham, Neale R. and Nyberg, Eric A. and Sillekens, Wim H. Magnesium for Crashworthy Components 463–466 Springer International Publishing 2016 Cham 10.1007/978-3-319-48099-2_75
[19] Biodegradable magnesium alloys for orthopaedic applications Biomaterials Translation 2021 2 3 214 10.12336/biomatertransl.2021.03.005
[20] Applications of magnesium alloys for aerospace: A review Journal of Magnesium and Alloys 2023 11 10 3609-3619 Magnesium and Its Alloys for Better Future - JMA 10th Anniversary https://doi.org/10.1016/j.jma.2023.09.015
[21] DAMASK – The Düsseldorf Advanced Material Simulation Kit for modeling multi-physics crystal plasticity, thermal, and damage phenomena from the single crystal up to the component scale Computational Materials Science 2019 158 420-478 https://doi.org/10.1016/j.commatsci.2018.04.030
[22] Expressing Crystallographic Textures through the Orientation Distribution Function: Conversion between Generalized Spherical Harmonic and Hyperspherical Harmonic Expansions Metallurgical and Materials Transactions A 2009 40 11 2590–2602 10.1007/s11661-009-9936-8
[23] Tuomo Nyyssönen and Azdiar A. Gazder and Ralf Hielscher and Frank Niessen Habit plane determination from reconstructed parent phase orientation maps Acta Materialia 2023 255 119035 https://doiȯrg/https://doiȯrg/10.1016/jȧctamat.2023.119035
[24] Machine learning pipeline for Structure–Property modeling in Mg-alloys using microstructure and texture descriptors Acta Materialia 2025 295 121132 https://doi.org/10.1016/j.actamat.2025.121132
[25] Inverse stochastic microstructure design Acta Materialia 2024 271 119877 https://doi.org/10.1016/j.actamat.2024.119877
[26] Active learning for the design of polycrystalline textures using conditional normalizing flows Acta Materialia 2025 284 120537 https://doi.org/10.1016/j.actamat.2024.120537
[27] Wei Xiong and Gregory B. Olson Cybermaterials: materials by design and accelerated insertion of materials npj Computational Materials 2016 2 1 15009 https://doiȯrg/10.1038/npjcompumats.2015.9
[28] Sample-efficient optimization in the latent space of deep generative models via weighted retraining Proceedings of the 34th International Conference on Neural Information Processing Systems 2020 NIPS '20 Curran Associates Inc.
[29] Local Latent Space Bayesian Optimization over Structured Inputs Advances in Neural Information Processing Systems 2022 S. Koyejo and S. Mohamed and A. Agarwal and D. Belgrave and K. Cho and A. Oh 35 34505–34518 Curran Associates, Inc.
[30] Accelerating Bayesian Optimization for Biological Sequence Design with Denoising Autoencoders Proceedings of the 39th International Conference on Machine Learning 2022 Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan 162 Proceedings of Machine Learning Research 20459–20478 PMLR
[31] High-dimensional Bayesian optimization with sparse axis-aligned subspaces Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence 2021 de Campos, Cassio and Maathuis, Marloes H. 161 Proceedings of Machine Learning Research 493–503 PMLR
[32] Dislocation Models of Crystal Grain Boundaries Phys. Rev. 1950 78 275–289 May 10.1103/PhysRev.78.275
[33] DREAM.3D: A Digital Representation Environment for the Analysis of Microstructure in 3D Integrating Materials and Manufacturing Innovation 2014 3 5 02 10.1186/2193-9772-3-5
[34] Christopher Yeung and Benjamin Pham and Ryan Tsai and Katherine T. Fountaine and Aaswath P. Raman DeepAdjoint: An All-in-One Photonic Inverse Design Framework Integrating Data-Driven Machine Learning with Optimization Algorithms ACS Photonics 2023 10 4 884–891 https://doiȯrg/10.1021/acsphotonics.2c00968
[35] TopologyGAN: Topology Optimization Using Generative Adversarial Networks Based on Physical Fields Over the Initial Domain 2020
[36] Antibody Design with Constrained Bayesian Optimization, Yimeng Zeng and Hunter Elliott and Phillip Maffettone and Peyton Greenside and Osbert Bastani and Jacob R. Gardner, ICLR 2024 Workshop on Generative and Experimental Perspectives for Biomolecular Design, 2024, https://openreviewṅet/forum?id=K5Sr6WSA4B
[37] Completely Derandomized Self-Adaptation in Evolution Strategies Evolutionary Computation 2001 9 2 159-195 06 10.1162/106365601750190398
[38] Navyanth Kusampudi and Martin Diehl Inverse design of dual-phase steel microstructures using generative machine learning model and Bayesian optimization International Journal of Plasticity 2023 171 103776 https://doiȯrg/https://doiȯrg/10.1016/j.ijplas.2023.103776
[39] {Ra{ß}loff Inverse design of spinodoid structures using Bayesian optimization Computational Mechanics 2026 77 1 275–296 https://doiȯrg/10.1007/s00466-024-02587-w
[40] Jaewan Park and Shashank Kushwaha and Junyan He and Seid Koric and Qibang Liu and Iwona Jasiuk and Diab Abueidda Nonlinear inverse design of mechanical multi-material metamaterials enabled by video denoising diffusion and structure identifier Engineering Applications of Artificial Intelligence 2026 172 114368 https://doiȯrg/https://doiȯrg/10.1016/jėngappai.2026.114368
[41] Li Zheng and Siddhant Kumar and Dennis M. Kochmann Algebraic language models for inverse design of metamaterials via diffusion transformers Nature Machine Intelligence 2026 8 4 628–640 https://doiȯrg/10.1038/s42256-026-01218-8
[42] Computational microstructure characterization and reconstruction: Review of the state-of-the-art techniques Progress in Materials Science 2018 95 1-41 https://doi.org/10.1016/j.pmatsci.2018.01.005
[43] Diffusion Models for Black-Box Optimization Proceedings of the 40th International Conference on Machine Learning 2023 Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan 202 Proceedings of Machine Learning Research 17842–17857 PMLR
[44] Training Diffusion Models with Reinforcement Learning ICML 2023 Workshop The Many Facets of Preference-Based Learning 2023
[45] Feedback Efficient Online Fine-Tuning of Diffusion Models Proceedings of the 41st International Conference on Machine Learning 2024 Salakhutdinov, Ruslan and Kolter, Zico and Heller, Katherine and Weller, Adrian and Oliver, Nuria and Scarlett, Jonathan and Berkenkamp, Felix 235 Proceedings of Machine Learning Research 48892–48918 PMLR
[46] Compositional Generative Inverse Design The Twelfth International Conference on Learning Representations 2024
[47] Physics-Informed Diffusion Models The Thirteenth International Conference on Learning Representations 2025
[48] Informed Machine Learning – A Taxonomy and Survey of Integrating Prior Knowledge into Learning Systems IEEE Transactions on Knowledge and Data Engineering 2023 35 1 614-633 10.1109/TKDE.2021.3079836
[49] Learning Transferable Visual Models From Natural Language Supervision Proceedings of the 38th International Conference on Machine Learning 2021 Marina Meila and Tong Zhang 139 Proceedings of Machine Learning Research 8748–8763 PMLR
[50] LAION-5B Advances in Neural Information Processing Systems 2022 35 25278–25294 Curran Associates, Inc.
[51] SDXL arXiv preprint arXiv:2307.01952 2023
[52] Scaling Rectified Flow Transformers for High-Resolution Image Synthesis arXiv preprint arXiv:2403.03206 2024
[53] Scalable Diffusion Models with Transformers Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2023 4195–4205
[54] Classifier-Free Diffusion Guidance arXiv preprint arXiv:2207.12598 2022
[55] Taming Transformers for High-Resolution Image Synthesis Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2021 12873–12883
[56] Flow Matching for Generative Modeling Proceedings of the International Conference on Learning Representations 2023
[57] VICReg Proceedings of the International Conference on Learning Representations 2022
[58] On the Continuity of Rotation Representations in Neural Networks 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2019 5738-5746 10.1109/CVPR.2019.00589
[59] Deep active optimization for complex systems Nature Computational Science 2025 5 9 801–812 10.1038/s43588-025-00858-x
[60] Active Subspaces: Emerging Ideas for Dimension Reduction in Parameter Studies Society for Industrial and Applied Mathematics 2015 2 SIAM Spotlights Philadelphia 10.1137/1.9781611973860
[61] Determinantal Point Processes for Machine Learning Foundations and Trends® in Machine Learning 2012 5 07 10.1561/2200000044
[62] GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium Advances in Neural Information Processing Systems 2017 30 6626–6637
[63] Zhou Wang and Eero P. Simoncelli and Alan Conrad Bovik Multiscale structural similarity for image quality assessment The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003 2003 2 1398–1402 Vol.2
[64] The Unreasonable Effectiveness of Deep Features as a Perceptual Metric CVPR 2018
[65] Cheng Wang and Xiaogui Wang and others A comparative study of plastic deformation behaviors of OFHC copper based on crystal plasticity models Journal of Materials Science 2021 56 8789–8814
[66] F. Wang and S. Sandlöbes and M. Diehl and L. Sharma and F. Roters and D. Raabe In situ observation of collective grain-scale mechanics in Mg and Mg–rare earth alloys Acta Materialia 2014 80 77–93
[67] Jingyu Zhang and Shurong Ding and Shiyu Du A damage-effect-involved phenomenological crystal plasticity model and computational methods for mechanical responses of FeCrAl alloys Materials Today Communications 2021 28 102595 https://doiȯrg/10.1016/jṁtcomm.2021.102595
[68] Physics-informed machine learning Nature Reviews Physics 2021 3 6 422–440 10.1038/s42254-021-00314-5
[69] Averaging Quaternions Journal of Guidance, Control, and Dynamics 2007 30 1193-1196 07 10.2514/1.28949