顯示具有 Pandas 標籤的文章。 顯示所有文章
顯示具有 Pandas 標籤的文章。 顯示所有文章

2021年3月15日 星期一

[ Python 常見問題 ] Pandas - How to find a minimum value in a pandas dataframe column ?

 Source From Here

Question
Examples of how to find a minimum value in a pandas dataframe column.

HowTo

Create a dataframe
Lets create for example a simple data frame:
  1. import pandas as pd  
  2.   
  3. data = {'Name':['Ben','Anna','Zoe','Tom','John','Steve','Becky','Bob'],   
  4.         'Age':[36,27,20,12,30,20,22,21]}  
  5.   
  6. df = pd.DataFrame(data)  
  7. df  
Output:
  1.     Name  Age  
  2. 0    Ben   36  
  3. 1   Anna   27  
  4. 2    Zoe   20  
  5. 3    Tom   12  
  6. 4   John   30  
  7. 5  Steve   20  
  8. 6  Becky   22  
  9. 7    Bob   21  
Find the min value in the column Age
To find the minimum value in the column Age, a solution is to use the pandas function min:
  1. df['Age'].min()  # Output: 12  
Find the index corresponding to the min value in the column Age
It is also possible to find the index corresponding to the min value in the column Age using the pandas function called idxmin:
  1. df['Age'].idxmin()  # Output: 3  
Then using the index above:
Output:
  1. Name    Tom  
  2. Age      12  
  3. Name: 3, dtype: object  
An example with multiple rows with a min value in the same column
Lets create a dataframe with two min values in the column Age:
  1. import pandas as pd  
  2.   
  3. data = {'Name':['Ben','Anna','Zoe','Tom','John','Steve','Becky','Bob'],   
  4.         'Age':[12,27,20,12,30,20,22,21]}  
  5.   
  6. df = pd.DataFrame(data)  
  7.   
  8. print(df)  
Output:
  1.     Name  Age  
  2. 0    Ben   12  
  3. 1   Anna   27  
  4. 2    Zoe   20  
  5. 3    Tom   12  
  6. 4   John   30  
  7. 5  Steve   20  
  8. 6  Becky   22  
  9. 7    Bob   21  
Then the function min:
  1. df['Age'].min()  # Output: 12  
however idxmin:
  1. df['Age'].idxmin()  # Output: 0  
To get rows with a min value in the column Age a solution is to do:
  1. df[ df['Age'] == df['Age'].min() ]  
Output:
  1.   Name  Age  
  2. 0  Ben   12  
  3. 3  Tom   12  
and to get the indexes:
  1. df[ df['Age'] == df['Age'].min() ].index  
which return:
  1. Int64Index([0, 3], dtype='int64')  


2021年3月13日 星期六

[ Python 常見問題 ] Pandas - Set value for particular cell in pandas DataFrame using index

 Source From Here

Question
I've created a Pandas DataFrame
  1. import pandas as pd  
  2.   
  3. df = pd.DataFrame(index=['A','B','C'], columns=['x','y'])  
  4. df  
Output:
  1.     x    y  
  2. A  NaN  NaN  
  3. B  NaN  NaN  
  4. C  NaN  NaN  
Then I want to assign value to particular cell, for example for row 'C' and column 'x'. I've expected to get such result:Any suggestions?
  1.     x    y  
  2. A  NaN  NaN  
  3. B  NaN  NaN  
  4. C  10  NaN  
Any suggestions?

HowTo
Going forward, the recommended method is .iat/.at. So the recommended approach is:


2021年3月10日 星期三

[ Python 常見問題 ] Pandas - Groupby: Count and mean combined

 Source From Here

Question
Working with PANDAS to try and summarise a dataframe as a count of certain categories, as well as the means sentiment score for these categories. There is table full of strings which have different sentiment scores, and I want to group each text source by saying how many posts they have, as well as the average sentiment of these posts.

My (simplified) dataframe looks like this:
  1. import pandas as pd  
  2. import numpy as np  
  3.   
  4. df = pd.DataFrame(data=[  
  5.     ['bar', 'some string', 0.13],  
  6.     ['foo', 'alt string',  -0.8],  
  7.     ['bar', 'another str',  0.7],  
  8.     ['foo', 'some text',   -0.2],  
  9.     ['foo', 'more text',   -0.5]],  
  10.     columns=['source', 'text', 'sent']  
  11. )  


My expected output will look like this:
  1. source    count     mean_sent  
  2. -----------------------------  
  3. foo       3         -0.5  
  4. bar       2         0.415  
HowTo
You can use groupby with aggregate:
  1. df.groupby('source') \  
  2.        .agg({'text':'size', 'sent':'mean'}) \  
  3.        .rename(columns={'text':'count','sent':'mean_sent'}) \  
  4.        .reset_index()  



[Git 常見問題] error: The following untracked working tree files would be overwritten by merge

  Source From  Here 方案1: // x -----删除忽略文件已经对 git 来说不识别的文件 // d -----删除未被添加到 git 的路径中的文件 // f -----强制运行 #   git clean -d -fx 方案2: 今天在服务器上  gi...