Pandas数据结构:Series详解
Pandas库提供了高级的数据结构和操作工具,极大简化了数据分析任务。它是基于NumPy构建的,广泛应用于数据处理领域。
本文主要介绍Pandas中的核心数据结构之一——Series。以下是关键内容:
1. Series概述
Series是一种类似于一维数组的对象,由一组数据及其对应的标签(索引)组成。默认情况下,如果未指定索引,会自动生成一个从0到N-1的整数型索引。
import pandas as pd
import numpy as np
data = pd.Series([10, 20, -30, 40])
print(data)
输出结果如下:
0 10
1 20
2 -30
3 40
dtype: int64
2. 自定义索引
可以为Series指定自定义索引,以便更灵活地标识数据点。
custom_data = pd.Series([10, 20, -30, 40], index=['a', 'b', 'c', 'd'])
print(custom_data)
输出结果:
a 10
b 20
c -30
d 40
dtype: int64
3. 通过字典创建Series
还可以使用字典直接创建Series对象,字典的键将作为索引。
dict_data = {'x': 100, 'y': 200, 'z': 300}
series_from_dict = pd.Series(dict_data)
print(series_from_dict)
输出:
x 100
y 200
z 300
dtype: int64
4. 处理缺失值
当某些索引不存在对应值时,Pandas会用NaN表示缺失值。
modified_index = ['x', 'y', 'w']
adjusted_series = pd.Series(dict_data, index=modified_index)
print(adjusted_series)
输出:
x 100.0
y 200.0
w NaN
dtype: float64
可以通过pd.isnull()和pd.notnull()检测缺失值。
print(pd.isnull(adjusted_series))
输出:
x False
y False
w True
dtype: bool
5. 索引与切片
支持通过索引访问或筛选数据。
print(custom_data['a'])
print(custom_data[['a', 'b', 'c']])
输出:
10
a 10
b 20
c -30
dtype: int64
6. 基本运算
支持与NumPy类似的数组运算,同时保留索引与值之间的关联。
filtered_data = custom_data[custom_data > 0]
scaled_data = custom_data * 2
exponential_data = np.exp(custom_data)
print(filtered_data)
print(scaled_data)
print(exponential_data)
输出:
a 10
b 20
d 40
dtype: int64
a 20
b 40
c -60
d 80
dtype: int64
a 22026.465795
b 485165195.41
c 0.000000
d 5.1847055286e+17
dtype: float64
7. 名称属性
Series对象及其索引均具有name属性,用于标识数据。
adjusted_series.name = "Sample Data"
adjusted_series.index.name = "Keys"
print(adjusted_series)
输出:
Keys
x 100.0
y 200.0
w NaN
Name: Sample Data, dtype: float64