Python for Everybody 中文版

Chapter 0 Base Material

| 关于   «  8. PY4E - Python for Everybody   ::   目录   ::   10. PY4E - 面向所有人的 Python  »

9. PY4E - Python for Everybody

切换导航

PY4E

第 1 章:简介 第 2 章:对象 第 3 章:条件判断 第 4 章:函数 第 5 章:迭代 第 6 章:字符串 第 7 章:文件 第 8 章:线性表 第 9 章:字典 第 10 章:元组 第 11 章:正则表达式 第 12 章:网络程序 第 13 章:Python 与 Web 服务 第 14 章:Python 对象 第 15 章:Python 与数据库 第 16 章:数据可视化

9.1. 线性表

9.1.1. 线性表是一个序列

与字符串类似,线性表 是值的序列。在字符串中,值是字符;在线性表中,它们可以是任何类型。线性表中的值称为 元素 或有时称为 项。

有几种创建新线性表的方法;最简单的是将元素放在方括号("[" 和 "]")中:

[10, 20, 30, 40]
['crunchy frog', 'ram bladder', 'lark vomit']

第一个示例是一个包含四个整数的线性表。第二个示例是一个包含三个字符串的线性表。线性表的元素不必是相同类型。下面的线性表包含一个字符串、一个浮点数、一个整数以及(瞧!)另一个线性表:

['spam', 2.0, 5, [10, 20]]

线性表内的线性表是 嵌套 的。

不包含元素的线性表称为空线性表;你可以使用空括号 `[]` 来创建一个。

正如你可能预期的那样,你可以将 list 值赋给变量:

>>> cheeses = ['Cheddar', 'Edam', 'Gouda']
>>> numbers = [17, 123]
>>> empty = []
>>> print(cheeses, numbers, empty)
['Cheddar', 'Edam', 'Gouda'] [17, 123] []

9.1.2. 线性表是可变的

访问线性表元素的语法与访问字符串字符的语法相同:方括号运算符。方括号内的表达式指定索引。记住索引从 0 开始:

>>> print(cheeses[0])
Cheddar

与字符串不同,线性表是可变的,因为你可以改变线性表中项目的顺序或重新分配线性表中的项目。当方括号运算符出现在赋值语句的左侧时,它标识了将被赋值的那个线性表元素。

>>> numbers = [17, 123]
>>> numbers[1] = 5
>>> print(numbers)
[17, 5]

numbers 列表的第三个元素(索引为 1),原本为 123,现在是 5。

你可以将线性表视为索引与元素之间的关系。这种关系称为 映射;每个索引“映射到”其中一个元素。

线性表索引的工作方式与字符串索引相同:

  • 任何整数表达式都可以用作索引。

  • 如果您尝试读取或写入一个不存在的元素,您会收到一个 IndexError。

  • 如果一个索引具有负值,则它从线性表的末尾向前计数。

之前的 `in` 运算符也适用于线性表。

>>> cheeses = ['Cheddar', 'Edam', 'Gouda']
>>> 'Edam' in cheeses
True
>>> 'Brie' in cheeses
False

9.1.3. 遍历线性表

遍历列表元素的最常用方法是使用 for 循环。其语法与字符串相同:

for cheese in cheeses:
    print(cheese)

这适用于仅需读取线性表元素的场景。但若要写入或更新元素,则需要索引。一种常见做法是将函数 range 与 len 组合:

for i in range(len(numbers)):
    numbers[i] = numbers[i] * 2

该循环遍历线性表并更新每个元素。len 返回线性表中的元素数量。range 返回从 0 到 n − 1 的索引列表,其中 n 是线性表的长度。每次循环时,i 获取下一个元素的索引。主体中的赋值语句使用 i 读取元素的旧值并分配新值。

对空列表的 for 循环永远不会执行其循环体:

for x in empty:
    print('This never happens.')

尽管一个线性表可以包含另一个线性表,但嵌套的线性表仍然计为一个元素。该线性表的长度为 4:

['spam', 1, ['Brie', 'Roquefort', 'Pol le Veq'], [1, 2, 3]]

9.1.4. 线性表操作

+ 运算符连接列表:

>>> a = [1, 2, 3]
>>> b = [4, 5, 6]
>>> c = a + b
>>> print(c)
[1, 2, 3, 4, 5, 6]

同样地,* 运算符将线性表重复给定次数:

>>> [0] * 4
[0, 0, 0, 0]
>>> [1, 2, 3] * 3
[1, 2, 3, 1, 2, 3, 1, 2, 3]

第一个示例重复四次。第二个示例重复线性表三次。

9.1.5. 线性表切片

切片运算符也适用于线性表:

>>> t = ['a', 'b', 'c', 'd', 'e', 'f']
>>> t[1:3]
['b', 'c']
>>> t[:4]
['a', 'b', 'c', 'd']
>>> t[3:]
['d', 'e', 'f']

如果省略第一个索引,切片从开头开始。如果省略第二个,切片直到末尾。因此,如果省略两者,切片是整个列表的副本。

>>> t[:]
['a', 'b', 'c', 'd', 'e', 'f']

由于线性表是可变的,因此在执行折叠、扭曲或破坏线性表的操作之前,通常很有用先制作一个副本。

赋值语句左侧的切片运算符可以更新多个元素:

>>> t = ['a', 'b', 'c', 'd', 'e', 'f']
>>> t[1:3] = ['x', 'y']
>>> print(t)
['a', 'x', 'y', 'd', 'e', 'f']

9.1.6. 线性表方法

Python 提供了用于操作线性表的方法。例如,append 向线性表的末尾添加一个新元素:

>>> t = ['a', 'b', 'c']
>>> t.append('d')
>>> print(t)
['a', 'b', 'c', 'd']

extend 接受一个线性表作为参数,并将所有元素追加到其中:

>>> t1 = ['a', 'b', 'c']
>>> t2 = ['d', 'e']
>>> t1.extend(t2)
>>> print(t1)
['a', 'b', 'c', 'd', 'e']

本示例未修改 t2。

sort 将线性表中的元素按从低到高的顺序排列:

>>> t = ['d', 'c', 'e', 'b', 'a']
>>> t.sort()
>>> print(t)
['a', 'b', 'c', 'd', 'e']

大多数线性表方法都是无返回值的;它们修改线性表并返回``None``。如果你不小心写了``t = t.sort()``,你会对结果感到失望。

9.1.7. 删除元素

有几种方法可以从列表中删除元素。如果你知道要删除的元素的索引,可以使用 pop:

>>> t = ['a', 'b', 'c']
>>> x = t.pop(1)
>>> print(t)
['a', 'c']
>>> print(x)
b

pop 修改线性表并返回被移除的元素。如果未提供索引,则删除并返回最后一个元素。

如果不需要被移除的值,可以使用 `del` 语句:

>>> t = ['a', 'b', 'c']
>>> del t[1]
>>> print(t)
['a', 'c']

如果您知道要移除的元素(但不知道其索引),可以使用 remove:

>>> t = ['a', 'b', 'c']
>>> t.remove('b')
>>> print(t)
['a', 'c']

remove 的返回值是 None。

若要移除多个元素,可以使用 `del` 配合切片索引:

>>> t = ['a', 'b', 'c', 'd', 'e', 'f']
>>> del t[1:5]
>>> print(t)
['a', 'f']

一如既往,切片选择直到第二个索引(但不包括该索引)之前的所有元素。

9.1.8. 线性表与函数

存在若干可在线性表上使用的内置函数,允许你无需编写自己的循环即可快速遍历线性表:

>>> nums = [3, 41, 12, 9, 74, 15]
>>> print(len(nums))
6
>>> print(max(nums))
74
>>> print(min(nums))
3
>>> print(sum(nums))
154
>>> print(sum(nums)/len(nums))
25

sum() 函数仅在列表元素为数字时有效。其他函数(如 max()、len() 等)可用于字符串列表及其他可比较的类型。

我们可以用线性表重写一个早期程序,该程序计算由用户输入的一组数字的平均值。

首先,计算平均值而不使用线性表的程序:

total = 0
count = 0
while (True):
    inp = input('Enter a number: ')
    if inp == 'done': break
    value = float(inp)
    total = total + value
    count = count + 1

average = total / count
print('Average:', average)

# Code: http://www.py4e.com/code3/avenum.py

在此程序中,我们拥有 count 和 total 两个变量,用于在反复提示用户输入数字时,分别记录用户输入的数字及其累加总和。

我们可以简单地记住用户输入的每个数字,并在最后使用内置函数计算总和与计数。

numlist = list()
while (True):
    inp = input('Enter a number: ')
    if inp == 'done': break
    value = float(inp)
    numlist.append(value)

average = sum(numlist) / len(numlist)
print('Average:', average)

# Code: http://www.py4e.com/code3/avelist.py

我们在循环开始前创建一个空线性表,然后每次遇到一个数字,就将其追加到线性表中。在程序结束时,我们简单地计算线性表中数字的总和,并将其除以线性表中数字的个数,从而得出平均值。

9.1.9. 线性表与字符串

字符串是字符的序列,而线性表是值的序列,但字符的线性表并不等同于字符串。要将字符串转换为字符的线性表,可以使用``list``:

>>> s = 'spam'
>>> t = list(s)
>>> print(t)
['s', 'p', 'a', 'm']

因为 list 是一个内置函数的名称,所以应避免将其用作变量名。我也避免使用字母 "l",因为它看起来太像数字 "1"。这就是为什么我使用 "t" 的原因。

list 函数将字符串拆分为单个字母。如果您想将字符串拆分为单词,可以使用 split 方法:

>>> s = 'pining for the fjords'
>>> t = s.split()
>>> print(t)
['pining', 'for', 'the', 'fjords']
>>> print(t[2])
the

一旦使用 `split` 将字符串拆分为单词的线性表后,可以使用索引运算符(方括号)来查看线性表中的特定单词。

您可以调用``split``并传入一个可选参数,称为 delimiter,用于指定用作单词边界的字符。以下示例使用连字符作为分隔符:

>>> s = 'spam-spam-spam'
>>> delimiter = '-'
>>> s.split(delimiter)
['spam', 'spam', 'spam']

join 是 split 的逆运算。它接收一个字符串列表并连接这些元素。join 是一个字符串方法,因此你需要在分隔符上调用它,并将列表作为参数传递:

>>> t = ['pining', 'for', 'the', 'fjords']
>>> delimiter = ' '
>>> delimiter.join(t)
'pining for the fjords'

在这种情况下,分隔符是一个空格字符,因此 join 在单词之间插入一个空格。若要连接不含空格的字符串,可以使用空字符串 "" 作为分隔符。

9.1.10. 解析行

通常当我们读取文件时,我们希望对行执行某些操作,而不仅仅是打印整行。我们通常希望找到“有趣的行”,然后 解析 该行以找到该行的某个有趣 部分。如果我们想从以“From”开头的那些行中打印出星期几,该怎么办?

From stephen.marquard@uct.ac.za Sat Jan  5 09:14:16 2008

split 方法在处理此类问题时非常有效。我们可以编写一个小程序来查找以 "From" 开头的行,split 这些行,然后打印出行中的第三个单词:

fhand = open('mbox-short.txt')
for line in fhand:
    line = line.rstrip()
    if not line.startswith('From '): continue
    words = line.split()
    print(words[2])

# Code: http://www.py4e.com/code3/search5.py

该程序产生以下输出:

Sat
Fri
Fri
Fri
...

后来,我们将学习越来越复杂的技术来挑选要处理的行,以及如何将这些行拆解,以找到我们正在寻找的确切信息。

9.1.11. 对象与值

如果我们执行这些赋值语句:

a = 'banana'
b = 'banana'

我们已知 a 和 b 都指向一个字符串,但我们不知道它们是否指向 同一个 字符串。存在两种可能的状态:

Variables and Objects

对象

在一种情况下,a 和 b 指代两个具有相同值的不同对象。在第二种情况下,它们指代同一个对象。

要检查两个变量是否引用同一个对象,可以使用 is 运算符。

>>> a = 'banana'
>>> b = 'banana'
>>> a is b
True

在此示例中,Python 仅创建了一个字符串对象,且 a 和 b 均指向它。

但当你创建两个线性表时,你会得到两个对象:

>>> a = [1, 2, 3]
>>> b = [1, 2, 3]
>>> a is b
False

在这种情况下,我们会说这两个线性表是 等价的,因为它们具有相同的元素,但不是 相同的,因为它们不是同一个对象。如果两个对象是相同的,那么它们也是等价的,但如果它们是等价的,则不一定相同。

迄今为止,我们一直将“对象”与“值”互换使用,但更精确的说法是:对象拥有一个值。如果你执行 a = [1,2,3],a 指的是一个列表对象,其值为特定元素序列。如果另一个列表具有相同的元素,我们会说它具有相同的值。

9.1.12. 别名

如果 a 指向一个对象并分配 b = a,则两个变量都指向同一个对象:

>>> a = [1, 2, 3]
>>> b = a
>>> b is a
True

变量与对象的关联称为 引用。在此示例中,有两个对同一对象的引用。

一个具有多个引用的对象具有多个名称,因此我们说该对象是 aliased。

如果对象是可变的,通过一个别名所做的更改会影响另一个:

>>> b[0] = 17
>>> print(a)
[17, 2, 3]

尽管这种行为可能有用,但它容易出错。一般来说,在处理可变对象时,避免别名更安全。

对于不可变对象(如字符串),别名化问题并不严重。在此示例中:

a = 'banana'
b = 'banana'

它几乎从不因 a 和 b 是否指向同一个字符串而产生差异。

9.1.13. 线性表参数

当您将线性表传递给函数时,该函数会获得该线性表的引用。如果函数修改了线性表参数,调用者将看到变化。例如,delete_head 从线性表中移除第一个元素:

def delete_head(t):
    del t[0]

以下是其用法:

>>> letters = ['a', 'b', 'c']
>>> delete_head(letters)
>>> print(letters)
['b', 'c']

参数 t 与变量 letters 是同一对象的别名。

区分修改线性表的操作与创建新线性表的操作非常重要。例如,append 方法修改线性表,但 + 操作符创建一个新的线性表:

>>> t1 = [1, 2]
>>> t2 = t1.append(3)
>>> print(t1)
[1, 2, 3]
>>> print(t2)
None

>>> t3 = t1 + [3]
>>> print(t3)
[1, 2, 3]
>>> t2 is t3
False

这种差异在编写旨在修改线性表(list)的函数时非常重要。例如,此函数 does not 删除线性表(list)的表头(delete_head):

def bad_delete_head(t):
    t = t[1:]              # WRONG!

切片运算符创建一个新线性表,赋值操作使 `t` 指向它,但这一切对作为参数传递的线性表没有任何影响。

另一种方法是编写一个函数来创建并返回一个新的线性表。例如,tail 返回线性表中除第一个元素外的所有元素:

def tail(t):
    return t[1:]

该函数不会修改原始线性表。其用法如下:

>>> letters = ['a', 'b', 'c']
>>> rest = tail(letters)
>>> print(rest)
['b', 'c']

练习 1:编写一个名为 ``chop`` 的函数,它接受一个列表并修改它,移除第一个和最后一个元素,并返回 ``None``。然后编写一个名为 ``middle`` 的函数,它接受一个列表并返回一个包含除第一个和最后一个元素之外所有元素的新列表。

9.1.14. 调试

随意使用列表(以及其他可变对象)可能导致长时间调试。以下是一些常见陷阱及避免方法:

  1. 别忘了大多数线性表方法会修改参数并返回 `None`。这与字符串方法相反,后者返回一个新字符串并保留原字符串不变。

如果你习惯这样编写字符串代码:

word = word.strip()

It is tempting to write list code like this:

t = t.sort()           # WRONG!

Because sort returns None, the next operation you perform with t is likely to fail.

Before using list methods and operators, you should read the documentation carefully and then test them in interactive mode. The methods and operators that lists share with other sequences (like strings) are documented at:

docs.python.org/library/stdtypes.html#common-sequence-operations

The methods and operators that only apply to mutable sequences are documented at:

docs.python.org/library/stdtypes.html#mutable-sequence-types

  1. 选择一个成语并坚持使用它。

列表的问题之一在于实现方式过多。例如,要从列表中删除一个元素,你可以使用 `pop`、`remove`、`del`,甚至切片赋值。

要添加一个元素,你可以使用 `append` 方法或 `+` 运算符。但别忘了这些是右侧的:

t.append(x)
t = t + [x]

And these are wrong:

t.append([x])          # WRONG!
t = t.append(x)        # WRONG!
t + [x]                # WRONG!
t = t + x              # WRONG!

Try out each of these examples in interactive mode to make sure you understand what they do. Notice that only the last one causes a runtime error; the other three are legal, but they do the wrong thing.

  1. 制作副本以避免别名。

如果您想使用像 sort 这样会修改参数的方法,但同时又需要保留原始线性表,您可以创建一个副本。

orig = t[:]
t.sort()

In this example you could also use the built-in function sorted, which returns a new, sorted list and leaves the original alone. But in that case you should avoid using sorted as a variable name!

  1. 列表、split 和文件

当我们读取和解析文件时,有许多机会会遇到可能导致程序崩溃的输入,因此在编写通过文件读取并寻找“大海捞针”的程序时,重新审视 guardian 模式是一个好主意。

重新审视我们的程序,该程序正在查找文件中的行以获取星期几:

来自 stephen.marquard@uct.ac.za Sat Jan  5 09:14:16 2008

由于我们将把该行拆分为单词,因此可以省去使用 startswith,只需查看该行的第一个单词即可确定我们是否对该行感兴趣。我们可以使用 continue 来跳过那些不以"From"作为第一个单词的行,方法如下:

fhand = open('mbox-short.txt')
for line in fhand:
    words = line.split()
    if words[0] != 'From' : continue
    print(words[2])

This looks much simpler and we don’t even need to do the rstrip to remove the newline at the end of the file. But is it better?

python search8.py
Sat
Traceback (most recent call last):
  File "search8.py", line 5, in <module>
    if words[0] != 'From' : continue
IndexError: list index out of range

It kind of works and we see the day from the first line (Sat), but then the program fails with a traceback error. What went wrong? What messed-up data caused our elegant, clever, and very Pythonic program to fail?

You could stare at it for a long time and puzzle through it or ask someone for help, but the quicker and smarter approach is to add a print statement. The best place to add the print statement is right before the line where the program failed and print out the data that seems to be causing the failure.

Now this approach may generate a lot of lines of output, but at least you will immediately have some clue as to the problem at hand. So we add a print of the variable words right before line five. We even add a prefix “Debug:” to the line so we can keep our regular output separate from our debug output.

for line in fhand:
    words = line.split()
    print('Debug:', words)
    if words[0] != 'From' : continue
    print(words[2])

When we run the program, a lot of output scrolls off the screen but at the end, we see our debug output and the traceback so we know what happened just before the traceback.

Debug: ['X-DSPAM-Confidence:', '0.8475']
Debug: ['X-DSPAM-Probability:', '0.0000']
Debug: []
Traceback (most recent call last):
  File "search9.py", line 6, in <module>
    if words[0] != 'From' : continue
IndexError: list index out of range

Each debug line is printing the list of words which we get when we split the line into words. When the program fails, the list of words is empty []. If we open the file in a text editor and look at the file, at that point it looks as follows:

X-DSPAM-Result: Innocent
X-DSPAM-Processed: Sat Jan  5 09:14:16 2008
X-DSPAM-Confidence: 0.8475
X-DSPAM-Probability: 0.0000

Details: http://source.sakaiproject.org/viewsvn/?view=rev&rev=39772

The error occurs when our program encounters a blank line! Of course there are “zero words” on a blank line. Why didn’t we think of that when we were writing the code? When the code looks for the first word (word[0]) to check to see if it matches “From”, we get an “index out of range” error.

This of course is the perfect place to add some guardian code to avoid checking the first word if the first word is not there. There are many ways to protect this code; we will choose to check the number of words we have before we look at the first word:

fhand = open('mbox-short.txt')
count = 0
for line in fhand:
    words = line.split()
    # print('Debug:', words)
    if len(words) == 0 : continue
    if words[0] != 'From' : continue
    print(words[2])

First we commented out the debug print statement instead of removing it, in case our modification fails and we need to debug again. Then we added a guardian statement that checks to see if we have zero words, and if so, we use continue to skip to the next line in the file.

We can think of the two continue statements as helping us refine the set of lines which are “interesting” to us and which we want to process some more. A line which has no words is “uninteresting” to us so we skip to the next line. A line which does not have “From” as its first word is uninteresting to us so we skip it.

The program as modified runs successfully, so perhaps it is correct. Our guardian statement does make sure that the words[0] will never fail, but perhaps it is not enough. When we are programming, we must always be thinking, “What might go wrong?”

练习 2:找出上面程序中哪一行仍然没有得到适当的保护。试着构造一个导致程序失败的文本文件,然后修改程序使得该行得到适当的保护,并测试它以确保能处理你的新文本文件。

练习 3:重写上述示例中的守护程序代码,不使用两个 ```if``` 语句。相反,使用包含 ```or``` 逻辑运算符的复合逻辑表达式,配合单个 ```if``` 语句。

9.1.15. 术语表

aliasing

两个或多个变量引用同一对象的情形。

delimiter

用于指示字符串分割位置的字符或字符串。

element

列表(或其他序列)中的一个值;也称为项。

equivalent

具有相同的值。

index

指示列表中某个元素的整数值。

identical

是同一个对象(这意味着等价)。

list

值的序列。

list traversal

按顺序访问列表中的每个元素。

nested list

作为另一个列表元素的列表。

object

变量可以引用的事物。对象具有类型和值。

reference

变量与其值之间的关联。

9.1.16. 练习

练习 4:找出文件中的所有唯一单词

莎士比亚在其作品中使用了超过 20,000 个单词。但您如何确定这一点?您如何生成莎士比亚使用的所有单词的列表?您会下载他的所有作品,阅读并手动追踪所有唯一单词吗?

让我们使用 Python 来实现这一目标。列出文件 romeo.txt 中存储的所有唯一单词,按字母顺序排序,该文件包含莎士比亚作品的一个子集。

为开始,下载该文件的副本 **www.py4e.com/code3/romeo.txt**。创建一个包含最终结果的唯一单词列表。编写一个程序打开文件 ``romeo.txt`` 并按行读取。对于每一行,使用 ``split`` 函数将行分割成单词列表。对于每个单词,检查该单词是否已在唯一单词列表中。如果单词不在唯一单词列表中,则将其添加到列表中。当程序完成时,按字母顺序对唯一单词列表进行排序并打印。

Enter file: romeo.txt
['Arise', 'But', 'It', 'Juliet', 'Who', 'already',
'and', 'breaks', 'east', 'envious', 'fair', 'grief',
'is', 'kill', 'light', 'moon', 'pale', 'sick', 'soft',
'sun', 'the', 'through', 'what', 'window',
'with', 'yonder']

练习 5:极简电子邮件客户端。

MBOX (mail box) 是一种流行的文件格式,用于存储和共享一组电子邮件。早期电子邮件服务器和桌面应用程序使用此格式。无需过多细节,MBOX 是一个文本文件,按顺序存储电子邮件。电子邮件由以 From 开头的特殊行分隔(注意空格)。重要的是,以 From: 开头的行(注意冒号)描述电子邮件本身,不作为分隔符。想象你编写了一个极简电子邮件应用程序,该程序列出用户收件箱中发件人的电子邮件,并计算电子邮件的数量。

编写一个程序,读取邮箱数据,当找到以 “From” 开头的行时,使用 ``split`` 函数将该行拆分为单词。我们关心是谁发送了消息,即 From 行上的第二个单词。

From stephen.marquard@uct.ac.za Sat Jan 5 09:14:16 2008

您将对 From 行进行解析,并打印出每一行 From 中的第二个单词,然后您还将统计 From(而非 From:)行的数量,并在最后打印出计数。这是一个很好的示例输出,其中省略了几行:

python fromcount.py
Enter a file name: mbox-short.txt
stephen.marquard@uct.ac.za
louis@media.berkeley.edu
zqian@umich.edu

[...some output removed...]

ray@media.berkeley.edu
cwen@iupui.edu
cwen@iupui.edu
cwen@iupui.edu
There were 27 lines in the file with From as the first word

练习 6:重写该程序,提示用户输入一个数字列表,当用户输入 “done” 时,在末尾打印出数字的最大值和最小值。将用户输入的数字存储在列表中,并在循环完成后使用 ``max()`` 和 ``min()`` 函数计算最大值和最小值。

Enter a number: 6
Enter a number: 2
Enter a number: 9
Enter a number: 3
Enter a number: 5
Enter a number: done
Maximum: 9.0
Minimum: 2.0

如果您在本书中发现错误,欢迎使用 Github 发送修正给我。

   «  8. PY4E - Python for Everybody   ::   目录   ::   10. PY4E - 面向所有人的 Python  »

关闭窗口