本文にスキップ
AI News HubLIVE
サイト内リライト2 分で読了

翻訳待ち:Feature Engineering in Scikit-Learn: A KDnuggets Cheat Sheet

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Once feature engineering lives inside a Pipeline, each step is fitted on training data only, and the model is scored what it actually earned. And that is the idea behind this new cheat sheet.

ソースKDnuggets著者: KDnuggets
翻訳待ち:Feature Engineering in Scikit-Learn: A KDnuggets Cheat Sheet
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Most of my early scikit-earn mistakes happened around a model rather than inside it. I would scale a column in one notebook cell, encode a category in another, fit the model somewhere further down, and then feel good about a cross-validation score that never survived contact with new data. Nothing was wrong with the estimators. The problem was that my preprocessing had already seen the validation fold before the model ever got there. What fixed it was not learning more transformers. It was learning where they belong. Once feature engineering lives inside a Pipeline, each step is fitted on training data only, and the model is scored what it actually earned. And that is the idea behind this new cheat sheet: everything on it is something you can drop into that chain. The pieces that generally get the most use are the boring structural ones. ColumnTransformer is how numeric and categorical columns get their own treatment without me splitting the frame by hand, and make_column_selector means I can pick columns by dtype instead of listing them, so a new column doesn't force me to edit the pipeline. SimpleImputer with add_indicator=True is a good combination to get into the habit of using, because the pattern of what was missing is often signal. On the categorical side, handle_unknown="ignore" in OneHotEncoder has saved me from more prediction-time crashes than I care to admit, and TargetEncoder is my default when cardinality gets high enough that one-hot encoding stops being reasonable. Two more earn their place for different reasons. set_output(transform="pandas") and get_feature_names_out() are the fastest way to see what your pipeline actually built, which matters when a ColumnTransformer and a PolynomialFeatures step may have turned twelve columns into an even hundred. And the syntax of GridSearchCV is the payoff: once preprocessing is inside the estimator, an imputation strategy becomes a hyperparameter like any other, and you can tune it alongside your regularization strength in one search. This cheat sheet collects those steps in one place, with the arguments that matter and the ones I keep forgetting.

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Once feature engineering lives inside a Pipeline, each step is fitted on training data only, and the model is scored what it actually earned. And that is the idea behind this new…

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。