Skip to content

Data Science with Python

Data Manipulation

DataFrame Function

The DataFrame() function is used to create a dataframe from a variety of sources.

Create a dataframe from a Python list

data = ['FTP', 'HTTP', 'SSH', 'HTTPS'] 
df = pd.DataFrame(data)

Create a dataframe and assign a column name

data = ['FTP', 'HTTP', 'SSH', 'HTTPS'] 
df = pd.DataFrame(data, columns=['Protocol'])

Create a dataframe with two columns, called Protocol and Port Number, from a Python nested-list

data = [['FTP', 21], ['HTTP', 80], ['SSH', 22], ['HTTPS', 443]] 
df = pd.DataFrame(data, columns=['Protocol', 'Port Number'])

The dtypes attribute of a dataframe holds the data types of the columns as shown below.

print(df.dtypes)

There are many ways of creating pandas dataframes, but most of the time, we will read data from a file, such as CSV or JSON, or download it from a source on the Internet.

Reading CSV Files

The read_csv() function reads a csv formatted file and returns a Pandas object.

Read a CSV file and create a dataframe from its content.

df = pd.read_csv('players.csv')

Print the first five rows of the dataframe

print(df.head())

Print the last five rows

print(df.tail())

Check the column names

print(df.columns)

Check the datatypes

print(df.dtypes)

usecols

The usecols parameter specifies the columns to be loaded into the dataframe. In the following example, we load only the Name, Number, Position, and College columns.

df = pd.read_csv('players.csv', usecols=['Name', 'Number', 'Position', 'College'])

Load all the columns but not ‘Height’ and ‘Weight’.

df = pd.read_csv('players.csv', usecols= lambda c : c not in ['Height', 'Weight'])

skip rows and nrows

The skiprows parameter allows us to skip specific rows while loading the csv file into the dataframe. In the following example, rows 1, 3, and 5 are not loaded.

df = pd.read_csv('players.csv', skiprows=[1, 3, 5])

Use the range() function to specify a range of rows to be excluded. In the following example, the first 100 rows of the file are excluded while loading it into the dataframe.

df = pd.read_csv('players.csv', skiprows=range(0, 100))

The nrows parameter allows loading a limited number of rows instead of all rows of a file. This argument is useful if the file is too large. For example, we load only the first 4 rows of the file as follows:

df = pd.read_csv('players.csv', nrows=4)

Visualization

Data Preprocessing

Clustering

Predictive Models (Classification Trees)

SVM (Classification)

NN (Regression)

Ensemble Learning