Frozen baseline untouched. Tracked file regenerated. Figure source `ManuscriptR1/images/BayesianSearchCV/` regenerated (R1 copy only). No FEM simulation or new expensive analysis was run; only existing result files were inspected and the figure source was re-exported.
Status key: **CLOSED** = fully incorporated; **PARTIALLY ADDRESSED** = improved with a reviewer sub-request still pending; **DEFERRED** = not actionable in this batch.
Evidence convention: `file:line` under the project root; model-output paths abbreviated as `Code/models/.../per_output_models_<W>W_B<B>_H<H>/it<k>/`.
---
## PART 0 — Verification of Batch 1
### A. Nested cross-validation wording — CORRECT
- Batch 1 states: inner folds for hyperparameter tuning/ranking, outer folds held out for generalization, all algorithms on the same outer splits; LOO for N≤20, repeated 5-fold × 3 for 21≤N≤80, 5-fold for N>80.
- Confirmed against `Code/src/width_optimization/surrogate_validation.py:40-90` (`RepeatedKFold(n_splits=5, n_repeats=3)`; inner `KFold` 3/4/5) and shared use in `ml_surrogate_train.py:816` and `rbf_surrogate_train.py:410`. Appendix A text (`~line 454`) corrected consistently.
-**Status: CLOSED** (no further edit required).
### B. 30-seeded DE runs — EXIST
- Production executes 30 independent runs per stage; raw per-run files exist with a header + 30 rows, e.g. `Code/reports/width_optimization/2W/ml_optimization/ml_2W_B29_H30_TFD_W100_it0_stage1_utilization_de_raw_runs.csv` (31 lines), `..._stage2_distortion_de_raw_runs.csv` (31 lines), `Code/reports/width_optimization/2W/rbf_optimization/rbf_2W_B29_H30_TFD_W100_it1_stage1_utilization_de_raw_runs.csv` (31 lines). Seeds `seed_base + run_idx`, base 42 (`de_utils.py:1422`).
- Therefore the Batch 1 statement "Each optimization stage is repeated with 30 independent runs using deterministic seeds derived from a fixed base seed" is retained (configuration + executed runs).
-**Status: CLOSED** (statement retained, supporting files identified). Integration of per-family reproducibility *numbers* is deferred (see missing analyses), because those results depend on the pending objective/results consolidation and must exclude F2.
---
## Detailed changes
### 1. Model-selection statistics moderated and contextualised
-**Original issue:** "clear hierarchy"; selection counts presented without caveat; "best overall performance"; reviewer R2#9 noted the tasks are correlated and selection frequency is not independent evidence.
-**Revised interpretation:** counts now given explicitly (SVR 71, GPR 15, GBR 10, XGBoost 2, MLP 2, RF 0); added the cross-check that the dispersion criterion rarely diverges from the single lowest-RMSE model (GPR lowest in 17 but selected in 15; GBR selected in 10 but lowest in 9); added that the tasks are not mutually independent and the counts are descriptive, not statistical proof of superiority; stated that the sensitivity to the 5% competitive threshold has not been assessed; SVR described as "most frequently selected" rather than "best overall".
-**Numerical/result evidence:**`ManuscriptR1/images/MLSurrogatesComparison/surrogate_selection_summary_by_variable_type.csv` (`selected_outputs`, `lowest_rmse_outputs`, `median_training_time_sec`); current per-fold selections in `Code/models/.../per_output_models_B29_H30/it0/nested_cv_fold_selections_...csv` (`selection_criterion = lowest_cv_rmse_dispersion_within_5pct_rmse_band`).
-**Original issue:** "kernel-based models are particularly well suited", "SVR provides the best compromise"; RBF training cost described via an unverified "<1 s"; no explicit statement that RBF and supervised models share the validation framework.
-**Revised interpretation:** kernel-based models "performed well under the adopted nested validation procedure"; explicit statement that RBF and supervised models use the **same outer splits and the same error metrics**, enabling a like-for-like comparison at each output; across outputs "neither strategy dominates"; RBF obtained with a much smaller hyperparameter search and lower training effort (specific "<1 second" claim removed).
-**Numerical/result evidence:**`Code/models/.../ml_models/.../nested_cv_outer_summary_*.csv` and `Code/models/.../rbf_models/.../nested_cv_outer_summary_*.csv` contain `outer_rmse, outer_mae, outer_r2, outer_std_abs_error, outer_max_abs_error` and bootstrap CI columns on the same splits.
-**Reviewer comment(s):** R2#8, R3#10.
-**Status: PARTIALLY ADDRESSED** — common framework stated and used; a consolidated normalized-RMSE/paired table is not precomputed (**DEFERRED**).
### 3. Small-sample robustness paragraph added
-**Section:** §5 Results (new paragraph after ~line 344).
-**Original issue:** reviewer concern about N=8 / N=16, model-selection optimism and uncertainty; the manuscript gave no explicit uncertainty discussion.
-**Revised interpretation:** new paragraph stating that results are constrained by small N; nested CV separates tuning from generalization; **bootstrap CIs** for RMSE/MAE are available; for N≤20 the outer loop is Leave-One-Out (single deterministic partition), with dispersion obtained by resampling the out-of-sample predictions; results bounded to the sampled design domain and analysed families; explicitly states **no independent FEM test set was retained** and that small-sample uncertainty remains a limitation.
-**Numerical/result evidence:** bootstrap CI columns in `nested_cv_outer_summary_*.csv` (`outer_rmse_ci_lower/upper`); deterministic outer LOO in `surrogate_validation.py:57-58`.
-**Reviewer comment(s):** R1#4, R2#7, R3#11.
-**Status: PARTIALLY ADDRESSED** — uncertainty is acknowledged and bounded; no repeated outer splitting or independent test set exists (**DEFERRED**).
### 4. Local FEM agreement distinguished from global predictive accuracy
-**Original issue:** R2#8/R3#10 warned that FEM error at one optimized candidate does not establish global surrogate superiority.
-**Revised interpretation:** added an explicit statement that `e_max` and `|e_J|` measure surrogate–FEM agreement **at the optimized candidate** (local consistency), not global accuracy; added a sentence stating these figures "are not a global ranking of predictive accuracy over the design domain", directing the reader to the outer-CV comparison.
-**Numerical/result evidence:** same as change 2 (outer-CV summaries are the global evidence; the FEM-validation table is local).
-**Reviewer comment(s):** R2#8, R3#10.
-**Status: CLOSED** (clarification; no new computation).
### 5. RBF-vs-ML FEM-agreement claims bounded
-**Section:** §5 Results (~line 400).
-**Original issue:** "RBF surrogates generally lead to lower objective-function discrepancies" presented without scope.
-**Revised interpretation:** prefixed with "For the set of optimized candidates analysed here"; retained the numerical averages (1.36 vs 5.04 for `|e_J|`; 1.79% vs 2.67% for `e_max`) as local-validation evidence.
### 6. Moderated "high predictive accuracy and robustness"
-**Section:** §5 Results (~line 417).
-**Original issue:** "Supervised models … provide high predictive accuracy and robustness" (robustness not demonstrated across seeds/splits).
-**Revised interpretation:** "provide competitive predictive accuracy under the adopted validation procedure".
-**Reviewer comment(s):** R2#8, R3#11.
-**Status: CLOSED**.
### 7. Bounded RBF surface-smoothness statement
-**Section:** §5 Results (~line 419).
-**Original issue:** "The response surfaces remain relatively smooth, supporting the suitability of RBF interpolation for this problem."
-**Revised interpretation:** bounded to "for the two-window families considered", consistent with the observed RBF performance for those datasets.
-**Status: CLOSED**.
### 8. Conclusions surrogate paragraph moderated
-**Section:** §6 Conclusions (~line 434).
-**Original issue:** "RBF interpolation proved to be even a better alternative … no hyperparameter search"; unbounded comparison.
-**Revised interpretation:** kernel-based models "most frequently selected"; when evaluated on the same outer splits, RBF "provided predictive performance comparable to the supervised models for the considered low-dimensional response surfaces" with "a much smaller hyperparameter search and substantially lower training effort"; for the optimized candidates RBF achieved surrogate–FEM agreement "comparable to, and in several cases better than" supervised ML; explicit domain-of-validity sentence ("limited to the analysed geometry families and to the sampled design domain") and larger uncertainty for the smallest datasets.
-**Reviewer comment(s):** R2#8, R3#10, R3#11.
-**Status: PARTIALLY ADDRESSED** — conclusion moderated and bounded; full like-for-like numeric consolidation still pending.
### 9. Highlight claim corrected
-**Section:** Highlights (~line 74).
-**Original issue:** "RBF interpolation matches supervised ML with lower retraining cost" — the exact claim R2#8 asked to retain only if a controlled comparison supports it.
-**Revised interpretation:** "RBF interpolation offers comparable accuracy to supervised ML at lower training cost".
-**Reviewer comment(s):** R2#8 (explicit).
-**Status: CLOSED**.
### 10. Abstract accuracy claim bounded
-**Section:** Abstract (~line 69).
-**Original issue:** "support vector regression and Gaussian process regression showing high predictive accuracy".
-**Revised interpretation:** "providing the most accurate predictions among the candidate models".
-**Original issue:** embedded text "Repeated K-Fold CV (K = 4, 5 times)" and "(40 evaluations)" were obsolete.
-**Revised interpretation:** source `ManuscriptR1/images/BayesianSearchCV/BayesianSearchCV.drawio` edited to "Repeated K-Fold CV (K = 5, 3 times)" and "(20, 30 or 40 evaluations)"; re-exported with draw.io 31.7.0 (`--crop`) to `BayesianSearchCV.pdf` (1 page, 1205×593 pt; original 1210×593 pt), plus `.svg`/`.png`. Visual style, layout and manuscript role preserved. The original `Manuscript/` figure was **not** modified.
-**Reviewer comment(s):** R2#7, R2#8, R3#11.
-**Status: CLOSED**.
---
## Analyses requested by reviewers but not found as completed results
1.**Competitive-threshold sensitivity (0%, 2.5%, 5%, 10%).** No such output exists (`find` for threshold/band CSVs returned only per-fold selection files). → Manuscript now states the sensitivity "has not been assessed"; **DEFERRED**.
2.**Normalized RMSE / paired per-task comparison across algorithms.** The nested-CV summaries hold `outer_rmse`, `outer_mae`, `outer_r2`, `outer_max_abs_error` per output, but no normalized aggregate or paired-task statistical comparison was found. → Like-for-like comparison stated at output level; aggregate table **DEFERRED**.
3.**Distribution/boxplots of per-algorithm normalized errors.** Not found. **DEFERRED**.
4.**Independent FEM test set not used in retraining.** Not found; explicitly acknowledged in the new paragraph. **DEFERRED**.
5.**Repeated outer resampling for N≤20.** Not implemented (outer is deterministic LOO); acknowledged. **DEFERRED**.
6.**Consolidated multi-seed DE variability table for the paper.** The 30-run raw/summary files exist, and `Code/reports/width_optimization/de_robustness_study/DE_ROBUSTNESS_SUMMARY.md` is a completed diagnostic. However, its representative examples include the B34 family (F2), and the per-family production results await the pending objective/results consolidation. → Not integrated into the manuscript; **DEFERRED** (no robustness numbers claimed).
7.**RBF training-time measurement file.** Not found (no timing/cost table for RBF). → The specific "<1 second per output" claim was removed; only qualitative lower training effort is retained. **DEFERRED**.
---
## Deliberately deferred (out of batch scope)
F2 final feasibility/results/iterations; objective-function reformulation and soft-target-vs-limit decision; distortion vs hysteretic-energy interpretation; FEM mesh sensitivity, boundary conditions, solver details; expanded experimental/FEM validation; system-level analyses; uncertainty studies requiring FEM; resilience; data/code availability; reviewer-response letter.