A foundational machine learning classification project using the Iris dataset to predict iris flower species from sepal and petal measurements.
This project was completed as part of an introductory machine learning class.
The objective was to build a machine learning system capable of classifying iris flowers into their respective species based on physical measurements of their sepals and petals.
The project follows a basic supervised machine learning workflow:
Data loading → Data inspection → Data cleaning → Exploratory analysis → Visualization → Encoding → Train/Test Split → Model Training → Model Evaluation
The project was framed around a botanical research institute seeking to automate iris species identification.
Instead of manually identifying flowers, the proposed machine learning system uses measurable flower characteristics to predict the species.
The target variable is:
Species
The input features are:
SepalLengthCmSepalWidthCmPetalLengthCmPetalWidthCm
The project uses the Iris dataset, containing measurements of iris flowers and their corresponding species.
The dataset includes three species:
- Iris-setosa
- Iris-versicolor
- Iris-virginica
The original dataset also contained an Id column. Since this was only a serial identifier and did not provide useful information for prediction, it was removed before modeling.
The dataset was inspected before modeling to understand its structure and quality.
The following checks were performed:
- Examined column names
- Reviewed sample and tail records
- Checked data types
- Checked for missing values
- Examined the distribution of the target classes
- Removed the
Idcolumn - Generated descriptive statistics
The three iris species were represented in the target variable, making this a multiclass supervised classification problem.
Several visualizations were used to understand relationships between the flower measurements and species.
A scatter plot was used to examine the relationship between:
- Sepal length
- Sepal width
with flower species represented separately.
A second scatter plot examined:
- Petal length
- Petal width
by species.
Histograms were also generated to inspect the distribution of the numerical variables.
Pearson correlation was calculated to examine relationships between the numerical variables.
A correlation heatmap was then created to make these relationships easier to interpret.
Because the dataset contained only a small number of features, all four measurements were retained for model training.
Machine learning algorithms generally require numerical input.
The Species column was therefore transformed from categorical labels into numerical values using Scikit-learn's LabelEncoder.
The dataset was divided into training and testing sets using:
- 70% training data
- 30% testing data
- Random state:
42
This allowed the models to be trained on one portion of the data and evaluated on previously unseen observations.
Two classification algorithms were trained and compared:
Gaussian Naive Bayes was used as one of the classification approaches.
The model achieved approximately 98% accuracy on the test dataset.
The confusion matrix showed that almost all flowers were correctly classified, with one Iris-versicolor observation classified as Iris-virginica.
KNN was also trained using the same features and training data.
The model achieved 100% accuracy on the test dataset for this particular train/test split.
The models were evaluated using:
- Accuracy score
- Confusion matrix
- Classification results
The confusion matrices were visualized using Scikit-learn's ConfusionMatrixDisplay.
The evaluation focused on how accurately the models classified previously unseen iris observations.
| Model | Test Accuracy |
|---|---|
| Gaussian Naive Bayes | 98% |
| K-Nearest Neighbors | 100% |
The results show that both algorithms performed strongly on this dataset.
The KNN model correctly classified all observations in the test set for the selected split.
- Python
- Pandas
- NumPy
- Matplotlib
- Seaborn
- Scikit-learn
- Jupyter Notebook
This project provided practical exposure to several fundamental machine learning concepts:
- Supervised learning
- Multiclass classification
- Feature and target identification
- Data cleaning
- Exploratory data analysis
- Data visualization
- Correlation analysis
- Label encoding
- Train/test splitting
- Model training
- Gaussian Naive Bayes
- K-Nearest Neighbors
- Accuracy evaluation
- Confusion matrices
This project demonstrated that sepal and petal measurements can be used to classify Iris flowers into three species using supervised machine learning. Both models performed strongly on the selected test set, with Gaussian Naive Bayes achieving approximately 98% accuracy and K-Nearest Neighbors achieving 100%.
The confusion matrix showed that the limited classification error occurred between Iris-versicolor and Iris-virginica, while Iris-setosa was classified correctly in the test results.
Beyond the model results, the project provided practical experience with the end-to-end machine learning workflow, from data inspection and exploratory analysis through feature preparation, model training, and evaluation.
Because this was an introductory project using a single train-test split, the reported accuracy should be viewed as a result for this experiment rather than a general measure of model performance.
iris-flower-classification/
│
├── ML_Iris Flower Project.ipynb
├── Iris.csv
├── iris-flower.jpg
└── README.md
Project Reflection
This project was an early introduction to applying machine learning to a structured dataset.
It helped build a foundation in the end-to-end classification workflow, from preparing data and exploring relationships to training models and evaluating predictions.
While the Iris dataset is a relatively simple educational dataset, the workflow introduced here forms part of the foundation for more complex machine learning and analytics projects.
Author
Ijeoma L. Anya
Operations Analytics | Data Analytics