manipulating a dataframe - pandas

nsx200 · (This post was last modified: Nov-13-2019, 03:23 PM by nsx200.)

Hello,

Whilst this is an assignment, I thought I would post in this forum page as it is pandas related...

I have a rubbish dataframe to use that lists sales amounts over a 6 month period. There are just two columns - a 'Country' column listing the names of countries that produced sales and a second column 'Sales Amount' that simply has an integer value for the number of sales for that country.

It is quite rubbish in that the countries appear multiple times (some more than others) and I have the task of creating a new dataframe with the unique country names as the index, but the second column must be a sum of the total number of sales for each country. So if England appears 3 times in the list with a 1, a 2 and a 3 being the integers in the three rows. I would need to return England in the first column with 6 in the second.

I have no issue getting the unique values, but my brain is stumped on producing the second column based on a total sum of each unique countries sales amounts.

So if I had the following basic dataframe...

df = pd.DataFrame{'England':1,'England':2,'England':3,'Wales':2,'Wales':7,'Wales':4}

Giving the output of...
England 1
England 2
England 3
Wales 2
Wales 7
Wales 4

I would need to return a new dataframe (not a pivot table or other structure - the assignment says it must be a dataframe) looking like this...

Country Sales Amount
England 6
Wales 13

Any pointers - not answers - they would be greatly appreciated.

I have tried to use a groupby() to populate the second column, but I just seem to get either 123 as the sales for England rather than 6, or I get 3 which is the number of times England appears in the Country column.

Apologies for the lack of code given, I am on my mobile and it is not easy to enter code with my very old and small phone screen.

Kind Regards

ThomasL · Nov-13-2019, 05:59 PM

Difficult to give you hints/pointers without giving the answer.
So, with .groupby() you were on the right track but didn´t take next step.

import pandas as pd

data = [['England', 1], ['England', 2], ['England', 3], ['Wales', 2], ['Wales', 7], ['Wales', 4]]
df = pd.DataFrame(data, columns=["country", "sales"])

new_df = df.groupby("country").sum()
new_df.reset_index(level=0, inplace=True)

print(new_df)
print(type(new_df))

nsx200 · Nov-14-2019, 10:38 AM

ThomasL

Thanks. Good to know I was nearly there initially.

I will perservere and look at the example you have given to make it work.

Thanks for your help

Possibly Related Threads…
Thread		Author	Replies	Views	Last Post
	Add NER output to pandas dataframe	dg3000	0	128	Apr-22-2024, 08:14 PM Last Post: dg3000
	HTML Decoder pandas dataframe column	mbrown009	3	1,061	Sep-29-2023, 05:56 PM Last Post: deanhystad
	Use pandas to obtain cartesian product between a dataframe of int and equations?	haihal	0	1,131	Jan-06-2023, 10:53 PM Last Post: haihal
	Pandas Dataframe Filtering based on rows	mvdlm	0	1,449	Apr-02-2022, 06:39 PM Last Post: mvdlm
	Pandas dataframe: calculate metrics by year	mcva	1	2,329	Mar-02-2022, 08:22 AM Last Post: mcva
	Pandas dataframe comparing	anto5	0	1,273	Jan-30-2022, 10:21 AM Last Post: anto5
	PANDAS: DataFrame \| Replace and others questions	moduki1	2	1,810	Jan-10-2022, 07:19 PM Last Post: moduki1
	PANDAS: DataFrame \| Saving the wrong value	moduki1	0	1,560	Jan-10-2022, 04:42 PM Last Post: moduki1
	update values in one dataframe based on another dataframe - Pandas	iliasb	2	9,324	Aug-14-2021, 12:38 PM Last Post: jefsummers
	empty row in pandas dataframe	rwahdan	3	2,461	Jun-22-2021, 07:57 PM Last Post: snippsat

manipulating a dataframe - pandas

User Panel Messages

Announcements