A machine learning project that predicts whether a college student will get placed based on their CGPA, IQ, and internship experience using Logistic Regression.
The dataset (college_student_placement_dataset.csv) contains the following features:
- College_ID: Unique identifier for each student
- IQ: Intelligence Quotient score
- CGPA: Cumulative Grade Point Average
- Internship_Experience: Whether the student has internship experience (Yes/No)
- Placement: Target variable - whether the student got placed (Yes/No)
Sample Data:
| College_ID | IQ | CGPA | Internship_Experience | Placement |
|---|---|---|---|---|
| CLG0030 | 107 | 6.28 | No | No |
| CLG0061 | 97 | 5.37 | No | No |
| CLG0036 | 109 | 5.83 | No | No |
| CLG0055 | 122 | 5.75 | Yes | No |
import pandas as pd
import numpy as np
df = pd.read_csv('college_student_placement_dataset.csv')
df.head()Output: First 5 rows of the dataset displayed.
# Convert categorical variables to numerical
df["Internship_Experience"] = df["Internship_Experience"].map({"Yes": 1, "No": 0})
df["Placement"] = df["Placement"].map({"Yes": 1, "No": 0})
# Sort by College_ID
df = df.sort_values(by="College_ID", ascending=True)
df.head()
len(df)Output:
- Categorical columns converted to binary (0/1)
- Dataset sorted by College_ID
- Total number of samples displayed
import matplotlib.pyplot as plt
plt.scatter(df["CGPA"], df["IQ"], c=df["Placement"])
plt.title("IQ vs CGPA")
plt.xlabel("CGPA")
plt.ylabel("IQ")Output: Scatter plot showing IQ vs CGPA colored by placement status (0=Not Placed, 1=Placed)
x = df.iloc[:, 1:3] # Features: IQ, CGPA
y = df.iloc[:, -1] # Target: PlacementNote: Internship_Experience was excluded from features (only IQ and CGPA used)
from sklearn.model_selection import train_test_split
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.1)Output: 90% training data, 10% test data
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
x_train = scaler.fit_transform(x_train)
x_test = scaler.transform(x_test)Output: Standardized features (mean=0, std=1) for both train and test sets
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression()
clf.fit(x_train, y_train)Output: Logistic Regression model trained successfully
y_pred = clf.predict(x_test)
y_test
from sklearn.metrics import accuracy_score
accuracy_score(y_test, y_pred)Output: Accuracy score on test set displayed
from mlxtend.plotting import plot_decision_regions
plot_decision_regions(x_train, y_train.values, clf=clf, legend=2)Output: Decision boundary plot showing how the model separates placed vs not placed students based on IQ and CGPA
import pickle
pickle.dump(clf, open('model.pkl', 'wb'))Output: Trained model saved as model.pkl for future use
- Install dependencies:
pip install pandas numpy matplotlib scikit-learn mlxtend- Run the notebook:
jupyter notebook main.ipynb- Use the saved model:
import pickle
import numpy as np
from sklearn.preprocessing import StandardScaler
# Load model
model = pickle.load(open('model.pkl', 'rb'))
# Prepare new data (IQ, CGPA)
new_data = np.array([[110, 7.5]]) # Example: IQ=110, CGPA=7.5
# Scale using same scaler (you'll need to fit on training data first)
# For production, save the scaler too!
scaler = StandardScaler()
# ... fit scaler on training data ...
new_data_scaled = scaler.transform(new_data)
# Predict
prediction = model.predict(new_data_scaled)
print("Placed" if prediction[0] == 1 else "Not Placed")| Metric | Value |
|---|---|
| Algorithm | Logistic Regression |
| Features Used | IQ, CGPA |
| Test Size | 10% |
| Preprocessing | StandardScaler |
| Model Saved | model.pkl |
- Include
Internship_Experienceas a feature - Try other algorithms (Random Forest, SVM, XGBoost)
- Perform hyperparameter tuning
- Add cross-validation
- Save the StandardScaler along with the model
- Create a simple web interface (Streamlit/Flask)
This project is for educational purposes.