# My first Machine Learning model

In this model, I attempt to analyze a set of training data and predict the survival of those who boarded the Titanic using Machine Learning’s Logistic Regression model.

Before we dive into the nitty-gritty of the model, let's talk briefly about data sets.

A dataset is a collection of data, usually presented in tabular form. It contains a set of observations or samples, each of which is described by a set of features or variables.

A dataset can come in many formats, such as a CSV file, an Excel spreadsheet, or a SQL table. It can also be stored in different types of databases, such as relational databases, NoSQL databases, or in-memory databases.

The data set we will be using today can be found here: [Titanic - Machine Learning from Disaster | Kaggle](https://www.kaggle.com/competitions/titanic/data)

**Now, let's get coding!**

These are the libraries we will be employing today. We'll learn about each one as we move forward -

```python
import numpy as numpy
import pandas as pandas
import matplotlib.pyplot as plot
import seaborn as seaborn
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
```

### Loading the data set:

Let's make a variable called titanic\_data and store the data in it. For doing so, we use the Pandas library.

*Pandas is a Python library for data manipulation and analysis. Pandas allows for easy handling of missing data and provides powerful data manipulation and cleaning capabilities, as well as data wrangling functionality.*

```python
titanic_data = pandas.read_csv('./data/train.csv')
```

To get a glimpse of the data we just stored in our data frame, print the following -

```python
# prints the first 5 rows of the dataframe
titanic_data.head()
# prints the total number of rows and Columns
titanic_data.shape
```

### Data pre-processing:

Data pre-processing is the process of preparing data for analysis by cleaning, transforming, and normalizing it. The goal of data pre-processing is to make the data suitable for use in machine learning models, as well as to improve the quality and accuracy of the results.

Let us first check the columns which have null values -  
`titanic_data.isnull().sum()`

When you run the above command, the result you see indicates the total number of null values in each column -

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1674228118600/d8708053-099f-4b9e-a221-e142e1885710.png align="left")

Since the **Cabin** column has the most number of null values and will not significantly contribute to our modeling, we can drop it.  
The **Age** column also has a considerable number of null values. A way of handling such data would be to substitute the mean value of all the ages that are present for the null values.  
Since **Embarked** has only 2 null values, we can substitute the mode (the value that appears most frequently) value.

```python
# drop the "Cabin" column because too many missing values
titanic_data = titanic_data.drop(columns='Cabin', axis=1)

# replacing the missing values in "Age" column with mean value
titanic_data['Age'].fillna(titanic_data['Age'].mean(), inplace=True)

# finding the mode value of "Embarked" column
print(titanic_data['Embarked'].mode()[0])

# replacing the missing values in "Embarked" column with mode value
titanic_data['Embarked'].fillna(titanic_data['Embarked'].mode()[0], inplace=True)

# check the number of missing values in each column again
titanic_data.isnull().sum()
```

After performing the above functions, we see the changes reflected in titanic\_data -

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1674228416769/39d311a8-c68f-4dc9-a467-4c0ebdd59400.png align="center")

### Data Visualisation:

Now that our data is processed, let's try to visualize it.  
Data visualization is the process of creating visual representations of data to communicate information effectively. The goal of data visualization is to make it easy to understand, explore and extract insights from large and complex datasets.

There are many ways to create visualizations, today we will be employing Matplotlib and Seaborn to visualize the Titanic data.

*Matplotlib is a plotting library for the Python programming language. It provides an object-oriented API for embedding plots into applications using general-purpose GUI toolkits.  
Seaborn is a data visualization library for Python, built on top of Matplotlib. It provides a high-level interface for drawing attractive and informative statistical graphics.*

```python
# set aspects of the visual theme for all matplotlib and seaborn plots.
seaborn.set()
# making a count plot for "Survived" column
seaborn.countplot(data=titanic_data, x='Survived')
```

The **countplot** function - "show the counts of observations in each categorical bin using bars."

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1674230158186/2ea5c578-3bf6-4e1e-aec8-d53cab0e0e6b.png align="center")

Here, we observe that the x-axis value corresponds to the "Survived" column of the data set and the y-axis gives us the corresponding count.

Similarly, a countplot for the column 'Sex' would look something like this -

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1674230262485/dce5bbc1-f13a-4cca-9ced-d3b360b7bf1c.png align="center")

Now let us try visualizing data against columns from the same data set.

```python
# number of survivors Gender wise
seaborn.countplot(data=titanic_data, x='Sex', hue='Survived', )
```

The result of this would look something like this -

![](https://cdn.hashnode.com/res/hashnode/image/upload/v1674230440025/22fa542e-c8bc-4b6f-96c4-e02aa8e3e0a1.png align="center")

Thus, the seaborn library helps us visualize the data we're handling. We can observe that even though there were more males present on the titanic, the number of female survivors was clearly more - which is a significant factor for the prediction model.

A significant step in the analyzing and preprocessing of data is to convert this data into categorical columns. Our prediction model will not care for string values such as male or female. Thus, let us encode our columns this way -

```python
titanic_data.replace({'Sex':{'male':0,'female':1}, 'Embarked':{'S':0,'C':1,'Q':2}}, inplace=True)
```

**Separating the features and targets**

In machine learning, features are the input variables used to make predictions, and the target is the output variable that is being predicted. It is important to separate features and targets in order to properly train and evaluate a model. The features are typically stored in a data matrix (X), while the target is stored in a separate vector (y). Once the data is separated into these two components, the model can be trained using the features to predict the target.

To do so we add the 'Survived' column to the Y vector and the columns apart from 'PassengerId', 'Name', 'Ticket', and 'Survived' in the X vector.

```python
# Separating features and targets
X = titanic_data.drop(columns = ['PassengerId','Name','Ticket','Survived'],axis=1)
Y = titanic_data['Survived']
```

**Test data vs Training data**

Now let us convert our features and targets into separate test and training data.

The training data will be used to train the model, while the test data will be used to evaluate the performance of the model.  
It's important to note that the model should not be trained on the test data, as this would lead to overfitting of the model and the performance would be artificially high.

```python
# Splitting the data into training and test sets

X_train, X_test, Y_train, Y_test = train_test_split(X,Y, test_size=0.2, random_state=3)

print(X.shape, X_train.shape, X_test.shape)
```

### Training the Model

Now that we have our training and testing sets, let's train the model.

**What is logistic regression?**

Logistic regression is a statistical method used for classification tasks. It is used to model the probability of a certain class or event occurring given a set of independent variables. Logistic regression is used for binary classification problems. ***This is why the logistic regression method is best suited for our binary (survive or not) classification problem statement.***

The **fit** function is used to train the model on the given data set.

```python
model = LogisticRegression(solver='lbfgs', max_iter=1000)
model.fit(X_train, Y_train)
```

Once our model is trained, let's find out the accuracy of the model achieved from the training set -

```python
# accuracy on training data
X_train_prediction = model.predict(X_train)

training_data_accuracy = accuracy_score(Y_train, X_train_prediction)

print('Accuracy score of training data : ', training_data_accuracy)
```

When we print the result, we see -

![](https://user-images.githubusercontent.com/43040456/213728266-8e392683-bbf7-42c0-8fee-4e0a85d53384.png align="left")

Now let's try to test our model -

```python
X_test_prediction = model.predict(X_test)

test_data_accuracy = accuracy_score(Y_test, X_test_prediction)

print('Accuracy score of test data : ', test_data_accuracy)
```

When we print the result, we see -

![](https://user-images.githubusercontent.com/43040456/213728314-ba478883-520b-4033-a735-7845523e7feb.png align="left")

Which was very close to our test data prediction. Thus our model is quite accurate as per the data we received.

### Checking the model for a random person:

Now, let’s check for a random person using random data from the unedited table from Kaggle.

```python
input = (3,0,35,0,0,8.05,0)  
# Note that this data excludes the Survived data, as it is to be determined from the model itself
```

Now let’s change these values to a NumPy array :

```python
input_as_numpy_array = numpy.asarray(input_data)
```

As our model was trained in different dimensions, we need to reshape this to our target dimensions.

```python
input_reshaped = input_as_numpy_array.reshape(1,-1)
```

Now, let’s predict using our model:

```typescript
prediction = model.predict(input_data_reshaped)
#print(prediction)
if prediction[0]==0:
    print("Dead")
if prediction[0]==1:
    print("Alive")
```

On running the code, we get the same result as that in the table.

Thus we can conclude that our model is performing well!
