AI・機械学習 / AI & Machine Learning

公開ケーススタディPublic case study

土壌VNIRスペクトルによる有機物含有量予測Soil Organic-Matter Content Prediction from VNIR Spectra

VNIRスペクトル回帰を、グループ分割・フォールド内前処理・予測区間まで含めて再設計した事例。

A redesigned VNIR spectral-regression workflow with grouped splits, fold-local preprocessing, and predictive intervals.

状態 / Status
修復済み・公開版は合成データRepaired / public synthetic workflow
言語 / Language
Python

技術スタック / Tech Stack

  • scikit-learn
  • Ridge regression
  • PLS regression
  • LASSO
  • Nested grouped cross-validation
  • Split conformal prediction

00

代表的な成果図 Selected Result Figure

土壌VNIRスペクトルによる有機物含有量予測の代表的な成果図 / Selected result figure for Soil Organic-Matter Content Prediction from VNIR Spectra
公開可能な検証結果から選んだ図です。 Selected from the validation evidence cleared for publication.

01

主なポイント Highlights

スペクトル前処理と特徴選択を交差検証フォールド内へ移し、情報漏洩を防止。

Moved spectral preprocessing and feature selection inside cross-validation folds to prevent leakage.

52サイトの公開合成フィクスチャで、グループ分割を含む全工程を再実行可能。

A public synthetic 52-site fixture supports a clean rerun of the full grouped workflow.

複数回帰器の比較に加え、split conformal法による予測不確実性を実装。

Implemented split-conformal predictive uncertainty alongside comparison of multiple regressors.

02

研究の流れ Research Flow

  1. 課題Problem

    高次元スペクトルから土壌有機物を予測する際のリークと過度に楽観的な評価を抑える。

    Reduce leakage and overly optimistic evaluation when predicting soil organic matter from high-dimensional spectra.

  2. 方法Method

    Ridge、PLS、LASSOを、サイト単位のネスト交差検証とフォールド内前処理で比較する。

    Compare Ridge, PLS, and LASSO with site-grouped nested cross-validation and fold-local preprocessing.

  3. 検証Validation

    公開合成データで再現性を確認し、グループ外評価と適合型予測区間を用いる。

    Validate reproducibility on public synthetic data using held-out groups and conformal prediction intervals.

  4. 成果Outcome

    実データを公開せずに、監査可能なモデリング手順と不確実性評価を提示。

    Demonstrates an auditable modeling and uncertainty workflow without publishing the real dataset.

03

限界と適用範囲 Limitations & Scope

04

公開範囲 Publication Boundary

公開ケーススタディPublic case study

実データは非公開。公開ページの図表と数値はすべて合成フィクスチャ由来で、方法と検証手順のみを紹介する。

The real data remain private. Every public figure and number comes from a synthetic fixture and is used only to present the method and validation workflow.

リポジトリRepositoryレポートReport再現手順Reproduceこの公開レベルでは利用できません / Not available at this publication tier