Skip to content

About

A foundational machine learning classification project using the Iris dataset to predict iris flower species from sepal and petal measurements.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

Repository files navigation

ML_Iris_Flower_Project

A foundational machine learning classification project using the Iris dataset to predict iris flower species from sepal and petal measurements.

iris-flower

Project Overview

This project was completed as part of an introductory machine learning class.

The objective was to build a machine learning system capable of classifying iris flowers into their respective species based on physical measurements of their sepals and petals.

The project follows a basic supervised machine learning workflow:

Data loading → Data inspection → Data cleaning → Exploratory analysis → Visualization → Encoding → Train/Test Split → Model Training → Model Evaluation

Business / Research Scenario

The project was framed around a botanical research institute seeking to automate iris species identification.

Instead of manually identifying flowers, the proposed machine learning system uses measurable flower characteristics to predict the species.

The target variable is:

  • Species

The input features are:

  • SepalLengthCm
  • SepalWidthCm
  • PetalLengthCm
  • PetalWidthCm

Dataset

The project uses the Iris dataset, containing measurements of iris flowers and their corresponding species.

The dataset includes three species:

  • Iris-setosa
  • Iris-versicolor
  • Iris-virginica

The original dataset also contained an Id column. Since this was only a serial identifier and did not provide useful information for prediction, it was removed before modeling.

Data Preparation

The dataset was inspected before modeling to understand its structure and quality.

The following checks were performed:

  • Examined column names
  • Reviewed sample and tail records
  • Checked data types
  • Checked for missing values
  • Examined the distribution of the target classes
  • Removed the Id column
  • Generated descriptive statistics

The three iris species were represented in the target variable, making this a multiclass supervised classification problem.

Exploratory Data Analysis

Several visualizations were used to understand relationships between the flower measurements and species.

Sepal Measurements

A scatter plot was used to examine the relationship between:

  • Sepal length
  • Sepal width

with flower species represented separately.

Petal Measurements

A second scatter plot examined:

  • Petal length
  • Petal width

by species.

Distribution

Histograms were also generated to inspect the distribution of the numerical variables.

Correlation Analysis

Pearson correlation was calculated to examine relationships between the numerical variables.

A correlation heatmap was then created to make these relationships easier to interpret.

Because the dataset contained only a small number of features, all four measurements were retained for model training.

Data Encoding

Machine learning algorithms generally require numerical input.

The Species column was therefore transformed from categorical labels into numerical values using Scikit-learn's LabelEncoder.

Train/Test Split

The dataset was divided into training and testing sets using:

  • 70% training data
  • 30% testing data
  • Random state: 42

This allowed the models to be trained on one portion of the data and evaluated on previously unseen observations.

Machine Learning Models

Two classification algorithms were trained and compared:

1. Gaussian Naive Bayes

Gaussian Naive Bayes was used as one of the classification approaches.

The model achieved approximately 98% accuracy on the test dataset.

The confusion matrix showed that almost all flowers were correctly classified, with one Iris-versicolor observation classified as Iris-virginica.

2. K-Nearest Neighbors (KNN)

KNN was also trained using the same features and training data.

The model achieved 100% accuracy on the test dataset for this particular train/test split.

Model Evaluation

The models were evaluated using:

  • Accuracy score
  • Confusion matrix
  • Classification results

The confusion matrices were visualized using Scikit-learn's ConfusionMatrixDisplay.

The evaluation focused on how accurately the models classified previously unseen iris observations.

Results

Model Test Accuracy
Gaussian Naive Bayes 98%
K-Nearest Neighbors 100%

The results show that both algorithms performed strongly on this dataset.

The KNN model correctly classified all observations in the test set for the selected split.

Tools & Technologies

  • Python
  • Pandas
  • NumPy
  • Matplotlib
  • Seaborn
  • Scikit-learn
  • Jupyter Notebook

Machine Learning Concepts Demonstrated

This project provided practical exposure to several fundamental machine learning concepts:

  • Supervised learning
  • Multiclass classification
  • Feature and target identification
  • Data cleaning
  • Exploratory data analysis
  • Data visualization
  • Correlation analysis
  • Label encoding
  • Train/test splitting
  • Model training
  • Gaussian Naive Bayes
  • K-Nearest Neighbors
  • Accuracy evaluation
  • Confusion matrices

Conclusion

This project demonstrated that sepal and petal measurements can be used to classify Iris flowers into three species using supervised machine learning. Both models performed strongly on the selected test set, with Gaussian Naive Bayes achieving approximately 98% accuracy and K-Nearest Neighbors achieving 100%.

The confusion matrix showed that the limited classification error occurred between Iris-versicolor and Iris-virginica, while Iris-setosa was classified correctly in the test results.

Beyond the model results, the project provided practical experience with the end-to-end machine learning workflow, from data inspection and exploratory analysis through feature preparation, model training, and evaluation.

Because this was an introductory project using a single train-test split, the reported accuracy should be viewed as a result for this experiment rather than a general measure of model performance.

Project Structure

iris-flower-classification/
│
├── ML_Iris Flower Project.ipynb
├── Iris.csv
├── iris-flower.jpg
└── README.md
Project Reflection

This project was an early introduction to applying machine learning to a structured dataset.

It helped build a foundation in the end-to-end classification workflow, from preparing data and exploring relationships to training models and evaluating predictions.

While the Iris dataset is a relatively simple educational dataset, the workflow introduced here forms part of the foundation for more complex machine learning and analytics projects.

Author

Ijeoma L. Anya

Operations Analytics | Data Analytics

About

A foundational machine learning classification project using the Iris dataset to predict iris flower species from sepal and petal measurements.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages