Showing posts with label dataframe. Show all posts
Showing posts with label dataframe. Show all posts

Monday, 29 January 2024

Pandas: replace string values in a DataFrame

Using 'replace()' method, we can perform string replacement operations in a dataframe

 

Example 1: Replace the values in a specific column.

df['City'] = df['City'].str.replace('Bangalore', 'Mumbai')

 

In the above example, the value 'Bangalore' in the 'City' column is replaced with 'Mumbai'. This is done by selecting the 'City' column using df['City'] and using the replace() method with the specified values to be replaced.

 

Example 2: Replace the values across dataframe by specifying a dictionary.

 

replace_dict = {'Krishna': 'Hari', 'Hyderabad': 'Delhi'}
df = df.replace(replace_dict)

 

replace_dict specifies the mapping values 'Krishna' to 'Hari' and 'Hyderabad' to 'Delhi'. The replace() method is then used on the entire DataFrame, replacing the values based on the dictionary mapping.

 

Find the below working application.

 

replace.py

 

import pandas as pd

# Create a sample DataFrame
data = {'Title': ['Krishna', 'Sailu', 'Joel', 'Chamu', 'Jitendra', "Krishna"],
        'Age': [34, 35, 234, 35, 52, 34],
        'City': ['Bangalore', 'Hyderabad', 'Hyderabad', 'Chennai', 'Bangalore', 'Chennai'],
        'Gender': ['Male', 'Female', 'Male', 'Female', 'Male', 'Male'],
        'Rating': [67, 43, 67, 100, 41, 89]}
df = pd.DataFrame(data)
print('Original DataFrame')
print(df)

df['City'] = df['City'].str.replace('Bangalore', 'Mumbai')
print('\nDataFrame after replacing City value Hyderabad to Mumbai\n', df)

# Replace values in multiple columns
replace_dict = {'Krishna': 'Hari', 'Hyderabad': 'Delhi'}
df = df.replace(replace_dict)
print('\nDataframe after replacing the dictionary of values\n', df)

Output

Original DataFrame
      Title  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      67
1     Sailu   35  Hyderabad  Female      43
2      Joel  234  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      41
5   Krishna   34    Chennai    Male      89

DataFrame after replacing City value Hyderabad to Mumbai
       Title  Age       City  Gender  Rating
0   Krishna   34     Mumbai    Male      67
1     Sailu   35  Hyderabad  Female      43
2      Joel  234  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52     Mumbai    Male      41
5   Krishna   34    Chennai    Male      89

Dataframe after replacing the dictionary of values
       Title  Age     City  Gender  Rating
0      Hari   34   Mumbai    Male      67
1     Sailu   35    Delhi  Female      43
2      Joel  234    Delhi    Male      67
3     Chamu   35  Chennai  Female     100
4  Jitendra   52   Mumbai    Male      41
5      Hari   34  Chennai    Male      89




 

Previous                                                 Next                                                 Home

Monday, 18 December 2023

Pandas: Filter rows of a DataFrame using where method

Using ‘where()’ method, we can filter the rows of a DataFrame.

I am using below data set to demonstrate the examples.

       Name  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      39
1     Sailu   35  Hyderabad  Female      43
2      Joel   29  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      41
5       Raj   34    Chennai    Male      89

Example 1: Get all the users who are from the city 'Bangalore'

bangalore_users = df.where(df['City'] == 'Bangalore')

‘bangalore_users’ point to the below data set.

        Name   Age       City Gender  Rating
0   Krishna  34.0  Bangalore   Male    39.0
1       NaN   NaN        NaN    NaN     NaN
2       NaN   NaN        NaN    NaN     NaN
3       NaN   NaN        NaN    NaN     NaN
4  Jitendra  52.0  Bangalore   Male    41.0
5       NaN   NaN        NaN    NaN     NaN

As you see above output, ‘bangalore_users’ contains NaN values for rows where the condition is not met, and it retains the original values where the condition is met.

 

You can also provide an alternative value to replace the NaN values by passing  the second argument to the where() method.

 

For example,

bangalore_users = df.where(df['City'] == 'Bangalore', 'not_matched')

In the above example, ‘bangalore_users’ point to below data set.

           Name          Age         City       Gender       Rating
0      Krishna           34    Bangalore         Male           39
1  not_matched  not_matched  not_matched  not_matched  not_matched
2  not_matched  not_matched  not_matched  not_matched  not_matched
3  not_matched  not_matched  not_matched  not_matched  not_matched
4     Jitendra           52    Bangalore         Male           41
5  not_matched  not_matched  not_matched  not_matched  not_matched

In this case, the rows that don't meet the condition are replaced with the string 'not_matched' in the resulting DataFrame.

 

Example 2: Get the rows whose age is 34 or Gender is 'Female'

users_age_is_34_or_female = df.where((df['Age'] == 34) | (df['Gender'] == 'Female'), 'not_matched')

 

'users_age_is_34_or_female' will point to below data set.

           Name          Age         City       Gender       Rating
0      Krishna           34    Bangalore         Male           39
1        Sailu           35    Hyderabad       Female           43
2  not_matched  not_matched  not_matched  not_matched  not_matched
3        Chamu           35      Chennai       Female          100
4  not_matched  not_matched  not_matched  not_matched  not_matched
5          Raj           34      Chennai         Male           89

Find the below working application.

 

where_method.py

import pandas as pd

# Create a sample DataFrame
data = {'Name': ['Krishna', 'Sailu', 'Joel', 'Chamu', 'Jitendra', "Raj"],
        'Age': [34, 35, 29, 35, 52, 34],
        'City': ['Bangalore', 'Hyderabad', 'Hyderabad', 'Chennai', 'Bangalore', 'Chennai'],
        'Gender': ['Male', 'Female', 'Male', 'Female', 'Male', 'Male'],
        'Rating': [39, 43, 67, 100, 41, 89]}
df = pd.DataFrame(data)
print('Original DataFrame')
print(df)

# Get the rows whose city is 'Bangalore'
bangalore_users = df.where(df['City'] == 'Bangalore')
print('\nbangalore_users\n', bangalore_users)

bangalore_users = df.where(df['City'] == 'Bangalore', 'not_matched')
print('\nbangalore_users\n', bangalore_users)

# Get the rows whose age is 34 or Gender is 'Female'
users_age_is_34_or_female = df.where((df['Age'] == 34) | (df['Gender'] == 'Female'), 'not_matched')
print('\nusers_age_is_35_or_female\n', users_age_is_34_or_female)

Output

Original DataFrame
       Name  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      39
1     Sailu   35  Hyderabad  Female      43
2      Joel   29  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      41
5       Raj   34    Chennai    Male      89

bangalore_users
        Name   Age       City Gender  Rating
0   Krishna  34.0  Bangalore   Male    39.0
1       NaN   NaN        NaN    NaN     NaN
2       NaN   NaN        NaN    NaN     NaN
3       NaN   NaN        NaN    NaN     NaN
4  Jitendra  52.0  Bangalore   Male    41.0
5       NaN   NaN        NaN    NaN     NaN

bangalore_users
           Name          Age         City       Gender       Rating
0      Krishna           34    Bangalore         Male           39
1  not_matched  not_matched  not_matched  not_matched  not_matched
2  not_matched  not_matched  not_matched  not_matched  not_matched
3  not_matched  not_matched  not_matched  not_matched  not_matched
4     Jitendra           52    Bangalore         Male           41
5  not_matched  not_matched  not_matched  not_matched  not_matched

users_age_is_35_or_female
           Name          Age         City       Gender       Rating
0      Krishna           34    Bangalore         Male           39
1        Sailu           35    Hyderabad       Female           43
2  not_matched  not_matched  not_matched  not_matched  not_matched
3        Chamu           35      Chennai       Female          100
4  not_matched  not_matched  not_matched  not_matched  not_matched
5          Raj           34      Chennai         Male           89

 

Previous                                                 Next                                                 Home

Pandas: nlargest: Get n largest values in a DataFrame

Using ‘nlargest’ method, we can get n largest values in a DataFrame.

 

I am going to use below data set to demonstrate the examples.

       Name  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      91
1     Sailu   35  Hyderabad  Female      76
2      Joel   29  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      87
5       Raj   34    Chennai    Male      89

 

Example 1: Get the three oldest people from the dataset.

three_oldest_emps = df.nlargest(3, columns='Age')

‘three_oldest_emps’ point to below data set.

        Name  Age       City  Gender  Rating
4  Jitendra   52  Bangalore    Male      87
1     Sailu   35  Hyderabad  Female      76
3     Chamu   35    Chennai  Female     100

Example 2: get the three oldest people by their rating from the dataset.

three_oldest_emps_by_rating = df.nlargest(3, columns=['Age', 'Rating'])

‘three_oldest_emps_by_rating’ point to below data set.

        Name  Age       City  Gender  Rating
4  Jitendra   52  Bangalore    Male      87
3     Chamu   35    Chennai  Female     100
1     Sailu   35  Hyderabad  Female      76

Find the below working application.

 

n_largest_rows.py

import pandas as pd

# Create a sample DataFrame
data = {'Name': ['Krishna', 'Sailu', 'Joel', 'Chamu', 'Jitendra', "Raj"],
        'Age': [34, 35, 29, 35, 52, 34],
        'City': ['Bangalore', 'Hyderabad', 'Hyderabad', 'Chennai', 'Bangalore', 'Chennai'],
        'Gender': ['Male', 'Female', 'Male', 'Female', 'Male', 'Male'],
        'Rating': [91, 76, 67, 100, 87, 89]}

df = pd.DataFrame(data)
print('Original DataFrame')
print(df)

three_oldest_emps = df.nlargest(3, columns='Age')
print('\nthree_oldest_emps : \n', three_oldest_emps)

three_oldest_emps_by_rating = df.nlargest(3, columns=['Age', 'Rating'])
print('\nthree_oldest_emps_by_rating : \n', three_oldest_emps_by_rating)

Output

Original DataFrame
       Name  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      91
1     Sailu   35  Hyderabad  Female      76
2      Joel   29  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      87
5       Raj   34    Chennai    Male      89

three_oldest_emps : 
        Name  Age       City  Gender  Rating
4  Jitendra   52  Bangalore    Male      87
1     Sailu   35  Hyderabad  Female      76
3     Chamu   35    Chennai  Female     100

three_oldest_emps_by_rating : 
        Name  Age       City  Gender  Rating
4  Jitendra   52  Bangalore    Male      87
3     Chamu   35    Chennai  Female     100
1     Sailu   35  Hyderabad  Female      76

 

Previous                                                 Next                                                 Home

Pandas: nsmallest: Get n smallest values in a DataFrame

Using ‘nsmallest’ method, we can get n smallest values in a DataFrame.

I am going to use below data set to demonstrate the examples.

       Name  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      91
1     Sailu   35  Hyderabad  Female      76
2      Joel   29  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      87
5       Raj   34    Chennai    Male      89

 

Example 1: Get the three youngest employees.

three_younger_emps = df.nsmallest(3, columns='Age')

 

‘three_younger_emps’ point to below data set.

       Name  Age       City Gender  Rating
2     Joel   29  Hyderabad   Male      67
0  Krishna   34  Bangalore   Male      91
5      Raj   34    Chennai   Male      89

Example 2: Get three youngest employees by their rating also.

three_younger_emps_by_rating = df.nsmallest(3, columns=['Age', 'Rating'])

‘three_younger_emps_by_rating’ point to below data set.

       Name  Age       City Gender  Rating
2     Joel   29  Hyderabad   Male      67
5      Raj   34    Chennai   Male      89
0  Krishna   34  Bangalore   Male      91

Find the below working application.

 

n_smallest_rows.py

import pandas as pd

# Create a sample DataFrame
data = {'Name': ['Krishna', 'Sailu', 'Joel', 'Chamu', 'Jitendra', "Raj"],
        'Age': [34, 35, 29, 35, 52, 34],
        'City': ['Bangalore', 'Hyderabad', 'Hyderabad', 'Chennai', 'Bangalore', 'Chennai'],
        'Gender': ['Male', 'Female', 'Male', 'Female', 'Male', 'Male'],
        'Rating': [91, 76, 67, 100, 87, 89]}

df = pd.DataFrame(data)
print('Original DataFrame')
print(df)

three_younger_emps = df.nsmallest(3, columns='Age')
print('\nthree_younger_emps : \n', three_younger_emps)

three_younger_emps_by_rating = df.nsmallest(3, columns=['Age', 'Rating'])
print('\nthree_younger_emps_by_rating : \n', three_younger_emps_by_rating)

Output

Original DataFrame
       Name  Age       City  Gender  Rating
0   Krishna   34  Bangalore    Male      91
1     Sailu   35  Hyderabad  Female      76
2      Joel   29  Hyderabad    Male      67
3     Chamu   35    Chennai  Female     100
4  Jitendra   52  Bangalore    Male      87
5       Raj   34    Chennai    Male      89

three_younger_emps : 
       Name  Age       City Gender  Rating
2     Joel   29  Hyderabad   Male      67
0  Krishna   34  Bangalore   Male      91
5      Raj   34    Chennai   Male      89

three_younger_emps_by_rating : 
       Name  Age       City Gender  Rating
2     Joel   29  Hyderabad   Male      67
5      Raj   34    Chennai   Male      89
0  Krishna   34  Bangalore   Male      91



Previous                                                 Next                                                 Home