From Data to Predictions- Understanding Scikit-learn in Python

Advertisements

If you are learning data science with Python, there is a good chance you will eventually encounter a library called Scikit-learn. At first glance, it may look like just another Python package filled with functions and classes. In practice, however, Scikit-learn represents something much more important. It provides one of the clearest bridges between the mathematical ideas behind machine learning and the practical process of building models from real data.

That distinction matters. Learning machine learning is not simply about knowing how to call a model and obtain a prediction. A competent data scientist needs to understand how data is prepared, how features are represented, how a model learns from those features, how its performance is evaluated, and how the entire process can be repeated reliably. Scikit-learn brings many of these steps together in a consistent framework, which is one of the main reasons it has become such an important part of the Python data-science ecosystem.

Scikit-learn is an open-source machine-learning library for Python designed primarily for traditional and classical machine-learning workflows. It provides implementations of many widely used algorithms for classification, regression, clustering, dimensionality reduction, preprocessing, model selection, and evaluation.

The library is built on top of the broader scientific Python ecosystem. In a typical data-science project, you might use NumPy for numerical operations, pandas for manipulating tabular data, Matplotlib or another visualization library for exploring patterns, and Scikit-learn for constructing and evaluating machine-learning models.

The real strength of Scikit-learn, though, is not merely the number of algorithms it contains. Its greatest advantage is consistency. Many different algorithms follow a similar interface, so once you understand the basic workflow, moving from one model to another becomes considerably easier.

For example, a model generally follows the same conceptual pattern: provide training data, allow the algorithm to learn from it, and then use the trained model to make predictions.

model.fit(X_train, y_train)
predictions = model.predict(X_test)

The simplicity of those two lines can be deceptive. Behind fit() lies the actual learning process, while predict() applies what the model learned to previously unseen observations. Understanding what happens between those two commands is where genuine machine-learning knowledge begins.

There is an important difference between writing machine-learning code and understanding machine learning.

A beginner can import a classifier, call fit(), obtain an accuracy score, and declare the project successful. A data scientist should immediately ask more difficult questions. Was the dataset representative? Was information from the test set accidentally used during training? Were the features appropriately scaled? Is accuracy actually the right metric? Is the model overfitting? Would a simpler model perform just as well?

Scikit-learn is valuable because it provides tools for investigating many of these questions.

Its functionality covers much more than individual algorithms. It includes preprocessing tools, train-test splitting, cross-validation, hyperparameter search, pipelines, feature selection, dimensionality reduction, and several evaluation metrics. In other words, it supports much of the experimental process through which a raw dataset becomes a machine-learning solution.

This is particularly useful for education. Instead of learning every algorithm as an isolated mathematical object, students can learn a general workflow and then examine how different algorithms behave within that workflow.

Imagine that you have a dataset containing information about houses and want to predict their prices. You might have features such as area, number of bedrooms, location-related variables, age of the property, and other characteristics.

The first conceptual distinction is between features and the target.

The features, commonly represented as X, are the information the model receives. The target, commonly represented as y, is what you want the model to predict.

In Python, a simplified representation might look like this:

X = data[["area", "bedrooms", "age"]]
y = data["price"]

The next step is usually to divide the available observations into training and testing data.

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

The purpose is fundamental. The model should learn from one portion of the data and then be evaluated on observations that were not used during training. Without this separation, it becomes difficult to determine whether the model has learned general patterns or has simply become very familiar with the examples it has already seen.

This idea of evaluating generalization is one of the foundations of machine learning.

Once the data has been divided, you can select an algorithm. Suppose you want to begin with linear regression.

from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
predictions = model.predict(X_test)

There is an elegant simplicity to this structure. You create an estimator, train it using fit(), and generate predictions using predict().

The important point is that Scikit-learn does not require you to implement the mathematical optimization procedure yourself. You do not need to manually derive the coefficients, calculate the loss function, or write an optimization routine simply to experiment with linear regression.

That abstraction is extremely useful. It allows the data scientist to concentrate on the larger problem: whether the chosen model and assumptions are appropriate for the data.

At the same time, abstraction should not become ignorance. If you use LinearRegression() without understanding what linear regression assumes and how it behaves, you may produce technically valid code that produces scientifically poor conclusions.

Not every problem involves predicting a numerical quantity such as price or temperature. Many practical problems require predicting a category.

You might want to determine whether an email is spam, whether a customer is likely to cancel a subscription, or whether a transaction should be classified as suspicious.

These are classification problems.

Scikit-learn provides several classification algorithms, including logistic regression, decision trees, support vector machines, nearest-neighbor methods, and ensemble techniques.

A simple example using logistic regression could look like this:

from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
predictions = model.predict(X_test)

The structure remains remarkably similar to the regression example. That consistency is one of Scikit-learn’s most useful characteristics.

Once you understand the estimator interface, the transition from one algorithm to another often involves changing the model rather than rewriting the entire machine-learning pipeline.

One of the most common mistakes among beginners is assuming that choosing a sophisticated algorithm is more important than preparing the data correctly.

It is not.

Machine-learning algorithms operate on numerical representations of information. Real-world datasets, however, are rarely ready for direct consumption. They may contain missing values, categorical variables, features with very different scales, or measurements expressed in incompatible formats.

Scikit-learn provides preprocessing utilities to address many of these issues.

For example, standardization can transform numerical features so that they are centered around zero and scaled according to their variability.

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Notice something subtle but extremely important here.

The scaler is fitted only on the training data. The test data is transformed using the parameters learned from the training set.

That distinction protects the evaluation process from a form of information leakage. If you allow information from the test set to influence preprocessing before evaluation, the resulting performance estimate can become overly optimistic.

This is a small technical detail with major consequences, and it illustrates why good data science depends as much on methodology as it does on algorithms.

As machine-learning projects become more complicated, manually performing every preprocessing step becomes increasingly error-prone.

This is where Scikit-learn’s Pipeline becomes particularly useful.

Instead of separately scaling the data and then training a model, you can combine the operations into a single workflow.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipeline = Pipeline([
("scaler", StandardScaler()),
("model", LogisticRegression())
])
pipeline.fit(X_train, y_train)
predictions = pipeline.predict(X_test)

The pipeline essentially says: first transform the data, then pass the transformed data to the model.

This approach has an important scientific advantage. It helps ensure that preprocessing is performed consistently during training, validation, and prediction. It also makes machine-learning experiments easier to reproduce and maintain.

For serious work, this is much more valuable than simply making the code shorter.

Advertisements

Training a model is only half of the problem. You also need to determine whether it performs well.

The appropriate evaluation metric depends on the problem.

For regression, common metrics include mean absolute error, mean squared error, and the coefficient of determination, commonly known as R².

For classification, accuracy can be useful in some situations, but it is far from universally appropriate. Precision, recall, F1 score, and other metrics can provide a more informative picture, particularly when the classes are imbalanced.

For example:

from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, predictions)
print(accuracy)

Suppose your classifier achieves 95 percent accuracy. That sounds impressive until you discover that 95 percent of the observations belong to one class. A model that simply predicts the majority class every time could achieve the same accuracy while being practically useless.

This is why evaluation should begin with the question, “What kind of mistake matters?”, rather than simply asking, “What is the accuracy?”

That change in perspective separates responsible model evaluation from superficial benchmarking.

One of the most important concepts you will encounter while learning Scikit-learn is overfitting.

A model can perform extremely well on the data used for training while performing poorly on new observations. In that situation, the model has effectively learned details that do not generalize.

Imagine a student who memorizes every answer in a practice exam but struggles when the questions are slightly changed. The student has learned the examples rather than the underlying concepts.

Machine-learning models can behave in much the same way.

Scikit-learn provides tools such as cross-validation that help researchers obtain a more reliable estimate of how a model is likely to perform on unseen data.

from sklearn.model_selection import cross_val_score
scores = cross_val_score(
model,
X,
y,
cv=5
)
print(scores.mean())

Instead of relying on a single train-test split, cross-validation repeatedly trains and evaluates the model across different partitions of the dataset.

The resulting estimate is generally more informative, particularly when the dataset is not especially large.

Another advantage of Scikit-learn is that it makes experimentation relatively straightforward.

Suppose you are working on a classification problem. Rather than assuming that one algorithm must be the best, you can compare several candidates.

You might begin with logistic regression, then investigate a decision tree, a random forest, or a support vector machine.

The purpose is not to collect algorithms for their own sake. It is to understand how different modeling assumptions interact with the structure of your data.

A linear model may perform remarkably well when the relationship between variables is approximately linear. A tree-based model may capture nonlinear relationships more naturally. A support vector machine may work particularly well in certain high-dimensional settings.

There is no universal algorithm that dominates every dataset.

The best model is usually discovered through disciplined experimentation rather than through loyalty to a particular technique.

Many Scikit-learn algorithms contain settings known as hyperparameters. These are parameters that control aspects of the learning algorithm but are not directly learned from the training observations in the same way as model coefficients.

For example, a decision tree can be constrained by its maximum depth. A random forest can be configured with a particular number of trees. A support vector machine has parameters that influence the shape of its decision boundary.

Choosing these values manually can become tedious. Scikit-learn provides tools such as GridSearchCV and RandomizedSearchCV to automate systematic experimentation with different configurations.

Conceptually, the process looks like this:

from sklearn.model_selection import GridSearchCV
search = GridSearchCV(
model,
param_grid,
cv=5,
scoring="accuracy"
)
search.fit(X_train, y_train)

The important lesson is that hyperparameter tuning should not be confused with simply making a model more complicated. The objective is to identify settings that provide strong generalization, not merely excellent performance on a particular training dataset.

Perhaps the most important thing to understand about Scikit-learn is that it should not be viewed as an algorithm encyclopedia.

Its real educational and practical value lies in the ecosystem surrounding the algorithms.

You can load data with pandas, inspect distributions and relationships through visualization, preprocess features with Scikit-learn, build a model, evaluate it using appropriate metrics, compare alternative approaches through cross-validation, and construct reproducible pipelines that combine multiple stages.

This gives Python data scientists a coherent framework for experimentation.

The library also teaches an important habit: thinking of machine learning as a process, not merely as a model.

A model is one component of a larger system that begins with a question and ends with a defensible conclusion.

Scikit-learn is exceptionally useful, but it is not the answer to every machine-learning problem.

It is particularly well suited to classical machine learning and structured or tabular data. It is not designed to replace specialized deep-learning frameworks when you are training large neural networks for computer vision, large language models, or other computationally intensive applications.

That distinction is worth understanding early.

If you are building a predictive model from a spreadsheet containing customer characteristics, financial measurements, business information, or scientific observations, Scikit-learn is often an excellent place to begin.

If you are training a deep neural network on millions of images or constructing a large transformer architecture, you will generally move into a different part of the Python ecosystem.

Knowing when not to use a tool is just as important as knowing how to use it.

The most effective approach is not to memorize every class available in the library.

Instead, build a small project from beginning to end.

Take a dataset, define the problem, inspect the variables, separate features from the target, divide the observations into training and testing sets, preprocess the data where necessary, train a simple baseline model, evaluate it properly, and then experiment with more sophisticated approaches.

For example, a beginner might build a model that predicts house prices, classify customer churn, identify handwritten digits, or estimate whether a customer will respond to an offer.

The particular dataset matters less than the process.

As you repeat that workflow, concepts that initially seem abstract begin to connect. Train-test splitting stops being a function you memorize and becomes a methodological safeguard. Cross-validation becomes a tool for estimating generalization. Pipelines become a way to prevent inconsistent preprocessing. Evaluation metrics become a reflection of the real-world consequences of prediction errors.

That is when Scikit-learn starts becoming more than a Python library. It becomes a framework through which you learn to think about data.

Once you are comfortable with the basic Scikit-learn workflow, the next step should not necessarily be learning dozens of additional algorithms.

Instead, deepen your understanding of the principles that make those algorithms useful.

Learn about feature engineering, data leakage, bias and variance, regularization, class imbalance, cross-validation, model interpretability, feature selection, dimensionality reduction, and hyperparameter optimization. These subjects will have a much greater impact on your ability to solve real problems than simply knowing the names of more algorithms.

Eventually, you should also become comfortable with the mathematics underlying the models you use. You do not need to become a theoretical mathematician, but understanding concepts such as probability, statistics, linear algebra, optimization, and statistical inference will make the behavior of machine-learning models far less mysterious.

This is ultimately the difference between using Scikit-learn and understanding it.

Scikit-learn is one of the best starting points for anyone who wants to move from Python programming into practical machine learning. Its consistent API makes complex algorithms accessible, while its preprocessing, evaluation, validation, and pipeline tools encourage a more disciplined approach to modeling.

But the most valuable lesson is not how to write model.fit().

It is learning what should happen before that line and what must happen after it.

Good data science begins with a meaningful question, continues with careful data preparation, uses appropriate statistical and machine-learning methods, and ends with an honest evaluation of what the model can and cannot tell us. Scikit-learn provides the machinery for much of that journey, but the quality of the final result still depends on the reasoning of the person using it.

If you are beginning your journey into data science, learning Scikit-learn is therefore not simply an exercise in learning another Python library. It is an opportunity to develop the workflow and analytical discipline that modern machine learning demands.

Advertisements

من البيانات إلى التنبؤات

في لغة بايثون Scikit-learn فهم مكتبة

Advertisements

إذا كنت تتعلم علم البيانات باستخدام لغة بايثون فمن المرجح جداً

Scikit-learn أن تصادف في مرحلة ما مكتبة تُدعى

فقد تبدو هذه المكتبة للوهلة الأولى مجرد حزمة برمجية أخرى في بايثون مليئة بالدوال والأصناف ولكنها في الواقع تمثل ما هو أهم من ذلك بكثير؛ فهي توفر أحد أوضح الجسور الرابطة بين المفاهيم الرياضية الكامنة وراء تعلم الآلة وبين العملية الفعالة لبناء النماذج استناداً إلى بيانات حقيقية

وهذا التمييز جوهري، فتعلم مجال تعلم الآلة لا يقتصر ببساطة على معرفة كيفية استدعاء نموذج ما والحصول على تنبؤ منه، إذ يحتاج عالم البيانات المتمكن إلى فهم كيفية إعداد البيانات وكيفية تمثيل السمات وكيفية تعلم النموذج من تلك السمات وكيفية تقييم أدائه وكيفية تكرار العملية بأكملها بشكل موثوق

العديد من هذه الخطوات Scikit-learn وتجمع مكتبة

في إطار عمل متسق ومنظم وهو أحد الأسباب الرئيسية التي جعلتها جزءاً بالغ الأهمية من منظومة علم البيانات في بايثون

هي مكتبة مفتوحة المصدر لتعلم الآلة Scikit-learn

مخصصة للغة بايثون وقد صُممت في المقام الأول لدعم مسارات العمل التقليدية والكلاسيكية في مجال تعلم الآلة، وهي توفر تطبيقات عملية للعديد من الخوارزميات واسعة الانتشار في مجالات التصنيف والانحدار والتجميع وتقليل الأبعاد والمعالجة المسبقة للبيانات واختيار النماذج وتقييمها تستند هذه المكتبة إلى منظومة بايثون العلمية الأوسع نطاقاً، ففي مشروع نموذجي لعلم البيانات

للعمليات الحسابية NumPy قد تستخدم مكتبة

لمعالجة البيانات الجدولية pandas ومكتبة

(أو أي مكتبة أخرى للرسوم البيانية) Matplotlib ومكتبة

لاستكشاف الأنماط

لبناء نماذج تعلم الآلة وتقييمها Scikit-learn بينما تستخدم

Scikit-learn ومع ذلك فإن القوة الحقيقية لمكتبة

لا تكمن فقط في عدد الخوارزميات التي تحتوي عليها بل تكمن ميزتها الكبرى في “الاتساق” إذ تتبع العديد من الخوارزميات المختلفة واجهة استخدام متشابهة مما يجعل الانتقال من نموذج إلى آخر أسهل بكثير بمجرد فهمك لمسار العمل الأساسي

على سبيل المثال تتبع النماذج عموماً النمط المفاهيمي ذاته: توفير بيانات التدريب والسماح للخوارزمية بالتعلم منها ثم استخدام النموذج المُدرَّب لإجراء التنبؤات

model.fit(X_train, y_train)
predictions = model.predict(X_test)

قد تكون بساطة هذين السطرين مضللة

تكمن عملية التعلم الفعلية `fit()` فخلف الدالة

بتطبيق ما تعلمه النموذج `predict()` بينما تقوم الدالة

على بيانات لم يسبق له رؤيتها، وهنا – أي في فهم ما يحدث بين هذين الأمرين – تبدأ المعرفة الحقيقية بتعلم الآلة

ثمة فرق جوهري بين كتابة شيفرة برمجية لتعلم الآلة وبين فهم تعلم الآلة ذاته

(classifier) يمكن للمبتدئ استيراد مُصنِّف

والحصول على درجة دقة `fit()` واستدعاء الدالة

ثم إعلان نجاح المشروع، أما عالم البيانات فينبغي عليه طرح أسئلة أكثر عمقاً وصعوبة على الفور: هل كانت مجموعة البيانات ممثلة للواقع؟ هل استُخدمت معلومات من مجموعة الاختبار عن طريق الخطأ أثناء التدريب؟

بشكل مناسب؟ (features) هل تمت مواءمة مقاييس السمات

هي المقياس الصحيح فعلاً؟ (accuracy) هل تُعد “الدقة”

؟ (overfitting)هل يعاني النموذج من فرط التخصيص

وهل يمكن لنموذج أبسط أن يحقق أداءً مماثلاً؟

في توفيرها أدوات Scikit-learn تكمن قيمة

لاستقصاء العديد من هذه الأسئلة

تتجاوز وظائف المكتبة مجرد توفير خوارزميات فردية فهي تشمل أدوات المعالجة المسبقة وتقسيم البيانات إلى مجموعتي تدريب واختبار والتحقق المتقاطع

(hyperparameters) والبحث عن المعلمات الفائقة

(pipelines) وخطوط المعالجة المتسلسلة

واختيار السمات وتقليل الأبعاد والعديد من مقاييس التقييم، بعبارة أخرى تدعم المكتبة جانباً كبيراً من العملية التجريبية التي تتحول فيها مجموعة بيانات خام إلى حل يعتمد على تعلم الآلة

وهذا مفيد بشكل خاص في المجال التعليمي، فبدلاً من تعلم كل خوارزمية ككيان رياضي منعزل يمكن للطلاب تعلم سير عمل عام ثم فحص كيفية أداء الخوارزميات المختلفة ضمن هذا المسار

تخيل أن لديك مجموعة بيانات تحتوي على معلومات حول منازل وتريد التنبؤ بأسعارها، فقد تتضمن البيانات سمات مثل المساحة وعدد غرف النوم ومتغيرات تتعلق بالموقع وعمر العقار وخصائص أخرى

“يتمثل التمييز المفاهيمي الأول في الفرق بين “السمات” و”الهدف

X السمات : التي يُرمز لها عادةً بالرمز

هي المعلومات التي يتلقاها النموذج

y أما الهدف : الذي يُرمز له عادةً بالرمز

فهو ما تريد من النموذج التنبؤ به

: في لغة بايثون قد يبدو التمثيل المبسط لذلك على النحو التالي

X = data[["area", "bedrooms", "age"]]
y = data["price"]

تتمثل الخطوة التالية عادةً في تقسيم الملاحظات المتاحة إلى بيانات للتدريب وبيانات للاختبار

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

يُعد هذا الغرض جوهرياً، إذ ينبغي للنموذج أن يتعلم من جزء من البيانات ثم يُقيَّم بناءً على مشاهدات لم تُستخدم أثناء مرحلة التدريب، فبدون هذا الفصل يصعب تحديد ما إذا كان النموذج قد تعلّم أنماطاً عامة أم أنه اكتفى بمجرد حفظ الأمثلة التي عُرضت عليه مسبقاً

وتُعد فكرة تقييم قدرة النموذج على التعميم هذه إحدى ركائز تعلّم الآلة

بمجرد تقسيم البيانات، يمكنك اختيار خوارزمية معينة، ولنفترض أنك ترغب في البدء بخوارزمية الانحدار الخطي

from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
predictions = model.predict(X_test)

“تتميز هذه البنية ببساطة أنيقة، إذ تقوم بإنشاء “مُقدِّر

`fit()` وتدريبه باستخدام الدالة

`predict()` ثم توليد التوقعات باستخدام الدالة

Scikit-learn وتكمن النقطة المهمة في أن مكتبة

لا تُلزمك بتنفيذ إجراءات التحسين الرياضي بنفسك، فلا داعي لاشتقاق المعاملات يدوياً أو حساب دالة الخسارة أو كتابة خوارزمية للتحسين لمجرد تجربة الانحدار الخطي

يُعد هذا المستوى من التجريد مفيداً للغاية، فهو يتيح لعالم البيانات التركيز على المشكلة الأكبر: وهي مدى ملاءمة النموذج والافتراضات المختارة للبيانات المتاحة وفي الوقت نفسه ينبغي ألا يتحول هذا التجريد إلى جهل

دون فهم الافتراضات `LinearRegression()` فإذا استخدمت

التي يقوم عليها الانحدار الخطي وآلية عمله فقد تكتب كوداً برمجياً صحيحاً من الناحية التقنية لكنه يؤدي إلى استنتاجات علمية ضعيفة أو غير دقيقة

Advertisements

لا تقتصر جميع المشكلات على التنبؤ بقيمة رقمية مثل السعر أو درجة الحرارة، فالعديد من المشكلات العملية تتطلب التنبؤ بفئة أو تصنيف معين على سبيل المثال: قد ترغب في تحديد ما إذا كانت رسالة البريد الإلكتروني

(spam) “رسالة مزعجة”

أو ما إذا كان من المحتمل أن يلغي العميل اشتراكه أو ما إذا كان ينبغي تصنيف معاملة مالية ما على أنها مشبوهة

(classification problems) تُعرف هذه الحالات بمشكلات التصنيف

العديد من خوارزميات التصنيف Scikit-learn توفر مكتبة

(logistic regression) بما في ذلك الانحدار اللوجستي

(decision trees) وأشجار القرار

(support vector machines) وآلات المتجهات الداعمة

(nearest-neighbor methods) وطرق الجار الأقرب

(ensemble techniques) وتقنيات النماذج المجمعة

: ويمكن أن يبدو المثال البسيط لاستخدام الانحدار اللوجستي على النحو التالي

from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
predictions = model.predict(X_test)

تظل البنية مشابهة بشكل لافت لمثال الانحدار ويُعد هذا الاتساق

فائدةً Scikit-learn إحدى أكثر ميزات مكتبة

(estimator) “فبمجرد استيعاب واجهة “المُقدِّر

غالباً ما يقتصر الانتقال من خوارزمية إلى أخرى على تغيير النموذج بدلاً من إعادة كتابة مسار عمل تعلم الآلة بالكامل

من أكثر الأخطاء شيوعاً بين المبتدئين افتراض أن اختيار خوارزمية متطورة أهم من إعداد البيانات بشكل صحيح، لكن الواقع خلاف ذلك تعمل خوارزميات تعلم الآلة بناءً على تمثيلات رقمية للمعلومات، في حين أن مجموعات البيانات الواقعية نادراً ما تكون جاهزة للاستخدام المباشر، إذ قد تحتوي على

(categorical variables) قيم مفقودة أو متغيرات فئوية

أو سمات ذات نطاقات قياس متفاوتة للغاية أو قياسات مُعبَّر عنها بتنسيقات غير متوافقة

أدوات للمعالجة الأولية Scikit-learn وتوفر مكتبة

تتيح معالجة العديد من هذه المشكلات

: فعلى سبيل المثال

(standardization) “يمكن لعملية “التوحيد القياسي

تحويل السمات الرقمية بحيث تتمركز حول الصفر وتُكيَّف مقاييسها وفقاً لدرجة تباينها

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

لاحظ هنا أمراً دقيقاً ولكنه بالغ الأهمية

بناءً على بيانات التدريب فقط (scaler) يتم ضبط أداة القياس

بينما تُحوَّل بيانات الاختبار

المستمدة من مجموعة التدريب (parameters) باستخدام المعلمات

يحمي هذا التمييز عملية التقييم من نوع من أنواع “تسرب المعلومات”، فإذا سمحت لمعلومات من مجموعة الاختبار بالتأثير على مرحلة المعالجة الأولية التي تسبق التقييم فقد يصبح تقدير الأداء الناتج متفائلاً بشكل مبالغ فيه

تُعد هذه تفصيلة تقنية صغيرة ذات عواقب كبيرة وهي توضح كيف تعتمد علوم البيانات الجيدة على المنهجية بقدر اعتمادها على الخوارزميات

مع ازدياد تعقيد مشاريع تعلم الآلة يصبح تنفيذ كل خطوة من خطوات المعالجة الأولية يدوياً أمراً عرضة للأخطاء بشكل متزايد وهنا تبرز أهمية وفائدة ميزة

Scikit-learn في مكتبة (Pipeline) “خط المعالجة”

(scaling) فبدلاً من إجراء عملية تحجيم البيانات

بشكل منفصل ثم تدريب النموذج يمكنك دمج هذه العمليات في مسار عمل واحد موحد

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipeline = Pipeline([
("scaler", StandardScaler()),
("model", LogisticRegression())
])
pipeline.fit(X_train, y_train)
predictions = pipeline.predict(X_test)

تتمثل آلية عمل خط المعالجة في الآتي: أولاً تُحوَّل البيانات ثم تُمرَّر البيانات المُحوَّلة إلى النموذج

ينطوي هذا النهج على ميزة علمية هامة، إذ يضمن إجراء المعالجة الأولية للبيانات بشكل متسق وموحد خلال مراحل التدريب والتحقق والتنبؤ، كما أنه يُسهِّل عملية إعادة إجراء تجارب التعلم الآلي وصيانتها وعند العمل على مشاريع جادة تكتسب هذه الميزة قيمة تفوق بكثير مجرد اختصار الكود البرمجي

إن تدريب النموذج لا يمثل سوى نصف المهمة، إذ يتعين عليك أيضاً تحديد مدى جودة أدائه

وتعتمد مقاييس التقييم المناسبة على طبيعة المشكلة المطروحة

: تشمل المقاييس الشائعة (regression) ففي مسائل الانحدار

(MAE) متوسط الخطأ المطلق

(MSE) ومتوسط مربع الخطأ

(R² ومعامل التحديد (المعروف اختصاراً بـ

(classification) أما في مسائل التصنيف

فقد تكون “الدقة” مفيدة في بعض الحالات لكنها ليست الخيار الأمثل في جميع الظروف، إذ يمكن لمقاييس أخرى

F1 (F1 score) مثل الدقة النوعية والاستدعاء ودرجة

أن تقدم صورة أكثر دقة وشمولاً لا سيما عندما تكون فئات البيانات غير متوازنة

: على سبيل المثال

from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, predictions)
print(accuracy)

لنفترض أن المُصنِّف ​​الخاص بك يحقق دقة بنسبة 95%. قد يبدو هذا أمراً مثيراً للإعجاب، إلى أن تكتشف أن 95% من البيانات (المشاهدات) تنتمي إلى فئة واحدة فقط، ففي هذه الحالة يمكن لنموذج يكتفي بتوقع فئة الأغلبية في كل مرة أن يحقق الدقة نفسها رغم أنه عديم الفائدة عملياً

ولهذا السبب يجب أن تبدأ عملية التقييم بطرح السؤال: “ما هو نوع الخطأ المهم؟” بدلاً من الاكتفاء بالسؤال: “ما هي نسبة الدقة؟”

إن تغيير وجهة النظر هذه هو ما يميز التقييم المسؤول للنموذج عن عمليات القياس المرجعي السطحية

(Overfitting يُعد “الفرط في التخصيص” (أو ما يُعرف بـ

Scikit-learn أحد أهم المفاهيم التي ستواجهها أثناء تعلم مكتبة

قد يُظهر النموذج أداءً ممتازاً للغاية عند التعامل مع البيانات المستخدمة في التدريب، بينما يكون أداؤه ضعيفاً عند التعامل مع بيانات جديدة، وفي هذه الحالة يكون النموذج قد تعلّم فعلياً تفاصيل لا يمكن تعميمها على حالات أخرى

تخيل طالباً يحفظ جميع إجابات امتحان تجريبي لكنه يواجه صعوبة إذا تغيرت الأسئلة قليلاً، فهذا الطالب قد تعلّم الأمثلة المحددة بدلاً من استيعاب المفاهيم الجوهرية الكامنة وراءها

ويمكن لنماذج تعلم الآلة أن تسلك المسار نفسه تماماً

أدوات Scikit-learn توفر مكتبة

(cross-validation) “مثل “التحقق المتقاطع

تساعد الباحثين في الحصول على تقدير أكثر موثوقية للأداء المتوقع للنموذج عند التعامل مع بيانات لم يسبق له رؤيتها

from sklearn.model_selection import cross_val_score
scores = cross_val_score(
model,
X,
y,
cv=5
)
print(scores.mean())

بدلاً من الاعتماد على تقسيم واحد للبيانات إلى مجموعتي التدريب والاختبار تقوم تقنية “التحقق المتقاطع” بتدريب النموذج وتقييمه بشكل متكرر عبر تقسيمات مختلفة لمجموعة البيانات

وعادةً ما يكون التقدير الناتج عن هذه العملية أكثر دلالة وفائدة لا سيما عندما لا تكون مجموعة البيانات كبيرة الحجم

Scikit-learn تتمثل ميزة أخرى لمكتبة

في أنها تجعل عملية التجريب بسيطة ومباشرة نسبياً

لنفترض أنك تعمل على حل مشكلة تصنيف، فبدلاً من افتراض أن خوارزمية معينة هي الأفضل حتماً يمكنك المقارنة بين عدة خوارزميات مرشحة

(logistic regression) قد تبدأ باستخدام الانحدار اللوجستي

ثم تختبر نماذج أخرى مثل شجرة القرار أو الغابة العشوائية أو آلة المتجهات الداعمة

والهدف هنا ليس مجرد تجميع الخوارزميات لذاتها بل فهم كيفية تفاعل افتراضات النمذجة المختلفة مع بنية بياناتك

قد يحقق النموذج الخطي أداءً ممتازاً عندما تكون العلاقة بين المتغيرات خطية تقريباً، بينما قد تتمكن النماذج القائمة على الأشجار من رصد العلاقات غير الخطية بشكل أكثر سلاسة وطبيعية، في حين قد تعمل آلة المتجهات الداعمة بكفاءة عالية في حالات معينة تتسم بأبعاد عالية للبيانات

لا توجد خوارزمية شاملة تتفوق على غيرها في جميع مجموعات البيانات

وعادةً ما يتم التوصل إلى النموذج الأفضل من خلال التجريب المنهجي والمنضبط وليس من خلال التمسك الأعمى بتقنية معينة

Scikit-learn تتضمن العديد من خوارزميات

إعدادات تُعرف باسم “المعلمات الفائقة” وهي عبارة عن معلمات تتحكم في جوانب معينة من خوارزمية التعلم، ولكنها لا تُستنبط أو تُتعلم مباشرة من بيانات التدريب بالطريقة نفسها التي تُحدد بها معاملات النموذج على سبيل المثال: يمكن تقييد شجرة القرار بتحديد حد أقصى لعمقها، ويمكن إعداد الغابة العشوائية بعدد محدد من الأشجار، كما تحتوي آلة المتجهات الداعمة على معلمات تؤثر في شكل

الخاص بها (decision boundary) “حد القرار”

قد تكون عملية اختيار هذه القيم يدوياً أمراً شاقاً ومملاً

Scikit-learn لذا توفر مكتبة

:أدوات مثل

`RandomizedSearchCV و `GridSearchCV`

لأتمتة عملية التجريب المنهجي باستخدام إعدادات مختلفة

:ومن الناحية المفاهيمية تبدو العملية على النحو التالي

from sklearn.model_selection import GridSearchCV
search = GridSearchCV(
model,
param_grid,
cv=5,
scoring="accuracy"
)
search.fit(X_train, y_train)

تتمثل العبرة المهمة في عدم الخلط بين ضبط المعلمات الفائقة وبين مجرد جعل النموذج أكثر تعقيداً، فالهدف هو تحديد الإعدادات التي تضمن قدرة قوية على التعميم وليس مجرد تحقيق أداء ممتاز على مجموعة بيانات تدريبية محددة

Scikit-learn ربما يكون أهم ما يجب فهمه حول مكتبة

هو أنه لا ينبغي النظر إليها باعتبارها موسوعة للخوارزميات فحسب

تكمن قيمتها التعليمية والعملية الحقيقية في المنظومة المتكاملة التي تحيط بهذه الخوارزميات

pandas إذ يمكنك تحميل البيانات باستخدام مكتبة

وفحص التوزيعات والعلاقات عبر أدوات التصور البياني

Scikit-learn ومعالجة السمات مسبقاً باستخدام

وبناء نموذج وتقييمه باستخدام المقاييس المناسبة ومقارنة الأساليب البديلة عبر التحقق المتقاطع

pipelines وإنشاء مسارات عمل

قابلة للتكرار تجمع بين مراحل متعددة

وهذا يوفر لعلماء البيانات الذين يستخدمون لغة بايثون إطار عمل متماسكاً لإجراء التجارب

كما تغرس هذه المكتبة عادةً مهمة: وهي التفكير في تعلم الآلة كعملية متكاملة وليس مجرد نموذج منفصل

فالنموذج هو جزء واحد من نظام أكبر يبدأ بطرح سؤال وينتهي باستنتاج يمكن تبريره ودعمه بالأدلة

مفيدة للغاية Scikit-learn تُعد مكتبة

لكنها ليست الحل الأمثل لكل مشكلة في مجال تعلم الآلة

فهي ملائمة بشكل خاص لأساليب تعلم الآلة التقليدية والبيانات المهيكلة أو الجدولية، وهي ليست مصممة لتحل محل أطر العمل المتخصصة في التعلم العميق عند تدريب شبكات عصبية ضخمة للرؤية الحاسوبية أو نماذج لغوية كبيرة أو غيرها من التطبيقات التي تتطلب موارد حاسوبية هائلة 

ومن المهم استيعاب هذا الفرق في مرحلة مبكرة إذا كنت بصدد بناء نموذج تنبؤي يعتمد على جدول بيانات يحتوي على خصائص العملاء أو مؤشرات مالية أو معلومات تجارية

Scikit-learn أو ملاحظات علمية فإن

غالباً ما تكون نقطة انطلاق ممتازة

أما إذا كنت تدرب شبكة عصبية عميقة على ملايين الصور أو تبني بنية ضخمة من نوع “المحولات” فستنتقل عادةً إلى جزء مختلف من منظومة بايثون البرمجية

إن معرفة متى لا يجب استخدام أداة معينة لا تقل أهمية عن معرفة كيفية استخدامها

النهج الأكثر فعالية ليس حفظ كل فئة متاحة في المكتبة

بدلاً من ذلك قم ببناء مشروع صغير من البداية وحتى النهاية، اختر مجموعة بيانات وحدد المشكلة وافحص المتغيرات وافصل الميزات عن الهدف وقسّم الملاحظات إلى مجموعتي تدريب واختبار وقم بمعالجة البيانات مسبقاً عند الضرورة ودرب نموذجاً أساسياً بسيطاً وقيمه بشكل صحيح ثم جرب أساليب أكثر تطوراً

على سبيل المثال: قد يبني مبتدئ نموذجاً يتنبأ بأسعار المنازل أو يصنف معدل فقدان العملاء أو يتعرف على الأرقام المكتوبة بخط اليد أو يقدر ما إذا كان العميل سيستجيب لعرض ما

لا تُعد مجموعة البيانات المحددة هي الأهم بل العملية نفسها

مع تكرار سير العمل هذا تبدأ المفاهيم التي بدت مجردة في البداية بالترابط، فلم يعد تقسيم التدريب والاختبار مجرد وظيفة تحفظها بل أصبح ضمانة منهجية، يصبح التحقق المتبادل أداة لتقدير التعميم وتصبح مسارات المعالجة وسيلة لمنع المعالجة المسبقة غير المتسقة وتصبح مقاييس التقييم انعكاساً للعواقب الواقعية لأخطاء التنبؤ

في أن يصبح أكثر من مجرد مكتبة بايثون Scikit-learn عندها يبدأ

يصبح إطار عمل تتعلم من خلاله التفكير في البيانات

Scikit-learn بمجرد أن تصبح ملماً بسير العمل الأساسي في مكتبة

فإن الخطوة التالية لا تقتصر بالضرورة على تعلم العشرات من الخوارزميات الإضافية

بدلاً من ذلك ينبغي عليك تعميق فهمك للمبادئ التي تجعل تلك الخوارزميات مفيدة

(feature engineering) تعرّف على مفاهيم مثل هندسة الميزات

(data leakage) وتسرب البيانات

(bias and variance) والانحياز والتباين

(regularization) والتنظيم

(class imbalance) وعدم توازن الفئات

(cross-validation) والتحقق المتقاطع

(model interpretability) وقابلية تفسير النموذج

(feature selection) واختيار الميزات

(dimensionality reduction) وتقليل الأبعاد

(hyperparameter optimization) وتحسين المعاملات الفائقة

سيكون لهذه الموضوعات تأثير أكبر بكثير على قدرتك على حل مشكلات واقعية مقارنة بمجرد معرفة أسماء المزيد من الخوارزميات

وفي النهاية ينبغي عليك أيضاً أن تصبح ملماً بالأسس الرياضية التي تقوم عليها النماذج التي تستخدمها، فلست بحاجة لأن تصبح عالم رياضيات نظرياً ولكن فهم مفاهيم مثل الاحتمالات والإحصاء والجبر الخطي والتحسين والاستدلال الإحصائي سيجعل سلوك نماذج تعلم الآلة أقل غموضاً بكثير

وهذا هو في النهاية الفرق

وفهمها حقاً Scikit-learn بين مجرد استخدام

واحدة من أفضل نقاط الانطلاق Scikit-learn تُعد

لأي شخص يرغب في الانتقال من برمجة بايثون إلى مجال تعلم الآلة العملي

المتسقة التي توفرها (API) فواجهة البرمجة

تجعل الخوارزميات المعقدة في متناول اليد، بينما تشجع أدواتها الخاصة بالمعالجة المسبقة والتقييم والتحقق وسلاسل العمليات على اتباع نهج أكثر انضباطاً في بناء النماذج

لكن الدرس الأكثر قيمة

`model.fit()` لا يكمن في كيفية كتابة الأمر

بل يكمن في تعلم ما يجب أن يحدث قبل هذا السطر وما يجب أن يحدث بعده يبدأ علم البيانات الجيد بسؤال ذي مغزى ويستمر بإعداد دقيق للبيانات ويستخدم أساليب إحصائية وأساليب تعلم آلة مناسبة وينتهي بتقييم صادق لما يمكن للنموذج أن يخبرنا به وما لا يمكنه ذلك

الأدوات اللازمة لجزء كبير من هذه الرحلة Scikit-learn وتوفر

لكن جودة النتيجة النهائية تظل معتمدة على التفكير والمنطق لدى الشخص الذي يستخدمها

لذا إذا كنت تبدأ رحلتك في علم البيانات

لا يمثل مجرد تمرين Scikit-learn فإن تعلم

لتعلم مكتبة أخرى في بايثون، بل هو فرصة لتطوير سير العمل والانضباط التحليلي الذي يتطلبه تعلم الآلة الحديث

Advertisements

Leave a Reply