AI・機械学習 / AI & Machine Learning

公開コードPublic code

データリークを防止したIris分類ベンチマークLeakage-Safe Iris Classification Benchmark

ネストした反復交差検証で、特徴量設計と複数分類器をデータリークなく比較するベンチマーク。

A leakage-safe benchmark comparing feature sets and classifiers with nested repeated cross-validation.

状態 / Status
再現済みベンチマークReproduced benchmark
言語 / Language
Python

技術スタック / Tech Stack

  • scikit-learn
  • Nested repeated stratified cross-validation
  • SVM
  • LDA
  • k-NN
  • Decision tree
  • Bootstrap and sign-flip inference

00

代表的な成果図 Selected Result Figure

データリークを防止したIris分類ベンチマークの代表的な成果図 / Selected result figure for Leakage-Safe Iris Classification Benchmark
公開可能な検証結果から選んだ図です。 Selected from the validation evidence cleared for publication.

01

主なポイント Highlights

外側10反復×5分割、内側3分割の設計でモデル選択と評価を分離。

Separated model selection from evaluation with 10 repeated 5-fold outer CV and 3-fold inner CV.

標準化やハイパーパラメータ調整を各訓練フォールド内に閉じ込め、リークを防止。

Kept scaling and hyperparameter tuning inside each training fold to prevent leakage.

SVMの2特徴量と4特徴量の差は統計的に決定的でなく、4特徴量LDAが記述的に最高精度。

The two- versus four-feature SVM difference was not statistically decisive, while four-feature LDA had the best descriptive accuracy.

02

研究の流れ Research Flow

  1. 課題Problem

    小規模分類で、特徴量追加の効果を過大評価せず複数モデルを公平に比較する。

    Compare classifiers fairly on a small dataset without overstating the benefit of adding features.

  2. 方法Method

    SVM、LDA、k近傍法、決定木をネストした反復層化交差検証で評価する。

    Evaluate SVM, LDA, k-NN, and decision trees using nested repeated stratified cross-validation.

  3. 検証Validation

    対応ブートストラップと符号反転検定で、同一分割上の特徴量差を評価する。

    Use paired bootstrap intervals and a sign-flip test to assess feature-set differences on identical splits.

  4. 成果Outcome

    再現可能な比較基盤を構築し、わずかな精度差と統計的不確実性を分けて報告。

    Produced a reproducible comparison that separates small accuracy differences from their statistical uncertainty.

03

限界と適用範囲 Limitations & Scope

04

公開範囲 Publication Boundary

公開コードPublic code

公開データ、再現コード、テスト、結果図を含む独立した公開リポジトリとして整備済み。

Prepared as a standalone public repository with open data, reproducible code, tests, and result figures.