feat: Add R1 Closed Changes — Batch 2 documentation and updates

parent 1a8000f5
# R1 Closed Changes — Batch 2
Editable file modified: `ManuscriptR1/ComparisonSurrogatesOptimizationBDSL_R1.tex`.
Frozen baseline untouched. Tracked file regenerated. Figure source `ManuscriptR1/images/BayesianSearchCV/` regenerated (R1 copy only). No FEM simulation or new expensive analysis was run; only existing result files were inspected and the figure source was re-exported.
Status key: **CLOSED** = fully incorporated; **PARTIALLY ADDRESSED** = improved with a reviewer sub-request still pending; **DEFERRED** = not actionable in this batch.
Evidence convention: `file:line` under the project root; model-output paths abbreviated as `Code/models/.../per_output_models_<W>W_B<B>_H<H>/it<k>/`.
---
## PART 0 — Verification of Batch 1
### A. Nested cross-validation wording — CORRECT
- Batch 1 states: inner folds for hyperparameter tuning/ranking, outer folds held out for generalization, all algorithms on the same outer splits; LOO for N≤20, repeated 5-fold × 3 for 21≤N≤80, 5-fold for N>80.
- Confirmed against `Code/src/width_optimization/surrogate_validation.py:40-90` (`RepeatedKFold(n_splits=5, n_repeats=3)`; inner `KFold` 3/4/5) and shared use in `ml_surrogate_train.py:816` and `rbf_surrogate_train.py:410`. Appendix A text (`~line 454`) corrected consistently.
- **Status: CLOSED** (no further edit required).
### B. 30-seeded DE runs — EXIST
- Production executes 30 independent runs per stage; raw per-run files exist with a header + 30 rows, e.g. `Code/reports/width_optimization/2W/ml_optimization/ml_2W_B29_H30_TFD_W100_it0_stage1_utilization_de_raw_runs.csv` (31 lines), `..._stage2_distortion_de_raw_runs.csv` (31 lines), `Code/reports/width_optimization/2W/rbf_optimization/rbf_2W_B29_H30_TFD_W100_it1_stage1_utilization_de_raw_runs.csv` (31 lines). Seeds `seed_base + run_idx`, base 42 (`de_utils.py:1422`).
- Therefore the Batch 1 statement "Each optimization stage is repeated with 30 independent runs using deterministic seeds derived from a fixed base seed" is retained (configuration + executed runs).
- **Status: CLOSED** (statement retained, supporting files identified). Integration of per-family reproducibility *numbers* is deferred (see missing analyses), because those results depend on the pending objective/results consolidation and must exclude F2.
---
## Detailed changes
### 1. Model-selection statistics moderated and contextualised
- **Section:** §5 Results (opening paragraph, ~line 335).
- **Original issue:** "clear hierarchy"; selection counts presented without caveat; "best overall performance"; reviewer R2#9 noted the tasks are correlated and selection frequency is not independent evidence.
- **Revised interpretation:** counts now given explicitly (SVR 71, GPR 15, GBR 10, XGBoost 2, MLP 2, RF 0); added the cross-check that the dispersion criterion rarely diverges from the single lowest-RMSE model (GPR lowest in 17 but selected in 15; GBR selected in 10 but lowest in 9); added that the tasks are not mutually independent and the counts are descriptive, not statistical proof of superiority; stated that the sensitivity to the 5% competitive threshold has not been assessed; SVR described as "most frequently selected" rather than "best overall".
- **Numerical/result evidence:** `ManuscriptR1/images/MLSurrogatesComparison/surrogate_selection_summary_by_variable_type.csv` (`selected_outputs`, `lowest_rmse_outputs`, `median_training_time_sec`); current per-fold selections in `Code/models/.../per_output_models_B29_H30/it0/nested_cv_fold_selections_...csv` (`selection_criterion = lowest_cv_rmse_dispersion_within_5pct_rmse_band`).
- **Reviewer comment(s):** R2#9 (primary), R2#7, R3#11.
- **Status: PARTIALLY ADDRESSED** — claims moderated; threshold-sensitivity analysis still absent (**DEFERRED**).
### 2. Like-for-like supervised-vs-RBF framing
- **Section:** §5 Results (kernel-model paragraph, ~line 344).
- **Original issue:** "kernel-based models are particularly well suited", "SVR provides the best compromise"; RBF training cost described via an unverified "<1 s"; no explicit statement that RBF and supervised models share the validation framework.
- **Revised interpretation:** kernel-based models "performed well under the adopted nested validation procedure"; explicit statement that RBF and supervised models use the **same outer splits and the same error metrics**, enabling a like-for-like comparison at each output; across outputs "neither strategy dominates"; RBF obtained with a much smaller hyperparameter search and lower training effort (specific "<1 second" claim removed).
- **Numerical/result evidence:** `Code/models/.../ml_models/.../nested_cv_outer_summary_*.csv` and `Code/models/.../rbf_models/.../nested_cv_outer_summary_*.csv` contain `outer_rmse, outer_mae, outer_r2, outer_std_abs_error, outer_max_abs_error` and bootstrap CI columns on the same splits.
- **Reviewer comment(s):** R2#8, R3#10.
- **Status: PARTIALLY ADDRESSED** — common framework stated and used; a consolidated normalized-RMSE/paired table is not precomputed (**DEFERRED**).
### 3. Small-sample robustness paragraph added
- **Section:** §5 Results (new paragraph after ~line 344).
- **Original issue:** reviewer concern about N=8 / N=16, model-selection optimism and uncertainty; the manuscript gave no explicit uncertainty discussion.
- **Revised interpretation:** new paragraph stating that results are constrained by small N; nested CV separates tuning from generalization; **bootstrap CIs** for RMSE/MAE are available; for N≤20 the outer loop is Leave-One-Out (single deterministic partition), with dispersion obtained by resampling the out-of-sample predictions; results bounded to the sampled design domain and analysed families; explicitly states **no independent FEM test set was retained** and that small-sample uncertainty remains a limitation.
- **Numerical/result evidence:** bootstrap CI columns in `nested_cv_outer_summary_*.csv` (`outer_rmse_ci_lower/upper`); deterministic outer LOO in `surrogate_validation.py:57-58`.
- **Reviewer comment(s):** R1#4, R2#7, R3#11.
- **Status: PARTIALLY ADDRESSED** — uncertainty is acknowledged and bounded; no repeated outer splitting or independent test set exists (**DEFERRED**).
### 4. Local FEM agreement distinguished from global predictive accuracy
- **Section:** §5 Results (FEM-validation paragraph ~line 346; RBF/ML comparison ~line 400).
- **Original issue:** R2#8/R3#10 warned that FEM error at one optimized candidate does not establish global surrogate superiority.
- **Revised interpretation:** added an explicit statement that `e_max` and `|e_J|` measure surrogate–FEM agreement **at the optimized candidate** (local consistency), not global accuracy; added a sentence stating these figures "are not a global ranking of predictive accuracy over the design domain", directing the reader to the outer-CV comparison.
- **Numerical/result evidence:** same as change 2 (outer-CV summaries are the global evidence; the FEM-validation table is local).
- **Reviewer comment(s):** R2#8, R3#10.
- **Status: CLOSED** (clarification; no new computation).
### 5. RBF-vs-ML FEM-agreement claims bounded
- **Section:** §5 Results (~line 400).
- **Original issue:** "RBF surrogates generally lead to lower objective-function discrepancies" presented without scope.
- **Revised interpretation:** prefixed with "For the set of optimized candidates analysed here"; retained the numerical averages (1.36 vs 5.04 for `|e_J|`; 1.79% vs 2.67% for `e_max`) as local-validation evidence.
- **Numerical/result evidence:** Table `tab:final_surrogate_comparison`; appendix tables.
- **Reviewer comment(s):** R2#8.
- **Status: CLOSED** (bounded wording).
### 6. Moderated "high predictive accuracy and robustness"
- **Section:** §5 Results (~line 417).
- **Original issue:** "Supervised models … provide high predictive accuracy and robustness" (robustness not demonstrated across seeds/splits).
- **Revised interpretation:** "provide competitive predictive accuracy under the adopted validation procedure".
- **Reviewer comment(s):** R2#8, R3#11.
- **Status: CLOSED**.
### 7. Bounded RBF surface-smoothness statement
- **Section:** §5 Results (~line 419).
- **Original issue:** "The response surfaces remain relatively smooth, supporting the suitability of RBF interpolation for this problem."
- **Revised interpretation:** bounded to "for the two-window families considered", consistent with the observed RBF performance for those datasets.
- **Status: CLOSED**.
### 8. Conclusions surrogate paragraph moderated
- **Section:** §6 Conclusions (~line 434).
- **Original issue:** "RBF interpolation proved to be even a better alternative … no hyperparameter search"; unbounded comparison.
- **Revised interpretation:** kernel-based models "most frequently selected"; when evaluated on the same outer splits, RBF "provided predictive performance comparable to the supervised models for the considered low-dimensional response surfaces" with "a much smaller hyperparameter search and substantially lower training effort"; for the optimized candidates RBF achieved surrogate–FEM agreement "comparable to, and in several cases better than" supervised ML; explicit domain-of-validity sentence ("limited to the analysed geometry families and to the sampled design domain") and larger uncertainty for the smallest datasets.
- **Reviewer comment(s):** R2#8, R3#10, R3#11.
- **Status: PARTIALLY ADDRESSED** — conclusion moderated and bounded; full like-for-like numeric consolidation still pending.
### 9. Highlight claim corrected
- **Section:** Highlights (~line 74).
- **Original issue:** "RBF interpolation matches supervised ML with lower retraining cost" — the exact claim R2#8 asked to retain only if a controlled comparison supports it.
- **Revised interpretation:** "RBF interpolation offers comparable accuracy to supervised ML at lower training cost".
- **Reviewer comment(s):** R2#8 (explicit).
- **Status: CLOSED**.
### 10. Abstract accuracy claim bounded
- **Section:** Abstract (~line 69).
- **Original issue:** "support vector regression and Gaussian process regression showing high predictive accuracy".
- **Revised interpretation:** "providing the most accurate predictions among the candidate models".
- **Reviewer comment(s):** R2#7, R4#1 (clarity/evidence).
- **Status: CLOSED**.
### 11. Figure `BayesianSearchCV` regenerated
- **Section:** §4.2, Figure `fig:BayesianSearchCV`.
- **Original issue:** embedded text "Repeated K-Fold CV (K = 4, 5 times)" and "(40 evaluations)" were obsolete.
- **Revised interpretation:** source `ManuscriptR1/images/BayesianSearchCV/BayesianSearchCV.drawio` edited to "Repeated K-Fold CV (K = 5, 3 times)" and "(20, 30 or 40 evaluations)"; re-exported with draw.io 31.7.0 (`--crop`) to `BayesianSearchCV.pdf` (1 page, 1205×593 pt; original 1210×593 pt), plus `.svg`/`.png`. Visual style, layout and manuscript role preserved. The original `Manuscript/` figure was **not** modified.
- **Reviewer comment(s):** R2#7, R2#8, R3#11.
- **Status: CLOSED**.
---
## Analyses requested by reviewers but not found as completed results
1. **Competitive-threshold sensitivity (0%, 2.5%, 5%, 10%).** No such output exists (`find` for threshold/band CSVs returned only per-fold selection files). → Manuscript now states the sensitivity "has not been assessed"; **DEFERRED**.
2. **Normalized RMSE / paired per-task comparison across algorithms.** The nested-CV summaries hold `outer_rmse`, `outer_mae`, `outer_r2`, `outer_max_abs_error` per output, but no normalized aggregate or paired-task statistical comparison was found. → Like-for-like comparison stated at output level; aggregate table **DEFERRED**.
3. **Distribution/boxplots of per-algorithm normalized errors.** Not found. **DEFERRED**.
4. **Independent FEM test set not used in retraining.** Not found; explicitly acknowledged in the new paragraph. **DEFERRED**.
5. **Repeated outer resampling for N≤20.** Not implemented (outer is deterministic LOO); acknowledged. **DEFERRED**.
6. **Consolidated multi-seed DE variability table for the paper.** The 30-run raw/summary files exist, and `Code/reports/width_optimization/de_robustness_study/DE_ROBUSTNESS_SUMMARY.md` is a completed diagnostic. However, its representative examples include the B34 family (F2), and the per-family production results await the pending objective/results consolidation. → Not integrated into the manuscript; **DEFERRED** (no robustness numbers claimed).
7. **RBF training-time measurement file.** Not found (no timing/cost table for RBF). → The specific "<1 second per output" claim was removed; only qualitative lower training effort is retained. **DEFERRED**.
---
## Deliberately deferred (out of batch scope)
F2 final feasibility/results/iterations; objective-function reformulation and soft-target-vs-limit decision; distortion vs hysteretic-energy interpretation; FEM mesh sensitivity, boundary conditions, solver details; expanded experimental/FEM validation; system-level analyses; uncertainty studies requiring FEM; resilience; data/code availability; reviewer-response letter.
---
## Supporting code and result files
- `Code/src/width_optimization/surrogate_validation.py` (CV and metrics), `ml_surrogate_train.py` (selection, Bayes n_iter, shared outer CV), `rbf_surrogate_train.py`, `rbf_model.py`.
- `Code/models/width_optimization/**/per_output_models_*/it*/nested_cv_outer_summary_*.csv`, `nested_cv_outer_predictions_*.csv`, `nested_cv_model_selection_frequency_*.csv`, `nested_cv_fold_selections_*.csv`, `final_model_selection_*.csv`.
- `Code/reports/width_optimization/**/**_de_raw_runs.csv` and `**_de_summary.csv` (30-run evidence).
- `Code/reports/width_optimization/de_robustness_study/DE_ROBUSTNESS_SUMMARY.md` (completed DE diagnostic; not integrated).
- `ManuscriptR1/images/MLSurrogatesComparison/surrogate_selection_summary_by_variable_type.csv` and `..._wide.csv`.
- `ManuscriptR1/images/BayesianSearchCV/BayesianSearchCV.drawio` (+ regenerated `.pdf`, `.svg`, `.png`).
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment