Friday, April 22, 2022

File processing

Most programming languages have built-in facilities for opening and processing files. In this next set of posts, I'll highlight Python's simple file input/output methods for handling text files.

Most of these facilities can be used for processing input from a variety of sources, such as reading from a URL. And we will do that at the end of this set of posts. For now, let's just get started with a simple set of exercises.

We'll be using the following file:


It's a plain text file, so it can be read by almost all programs that read files. This particular file has 14 lines. Using the VI editor, you can get a line count by typing the following

:set number

or

:set nu

Make sure that you're not in INSERT mode (i.e., the word INSERT does not appear at the bottom left corner of the editor window).

Now onto programming.

Our file is named test.txt. In python, we assign a file variable using the open() function. We will then use this file variable to manipulate the contents of the file, such as reading or writing to the file.

Start your python editor and type in the following. Make sure that you are in the same folder that your file is located.


The file is opened using the following code:

f = open('test.txt', 'r')

f is the file handle, or variable, that we will use to manipulate the file contents.

open() is the function used to open the file. open() takes two arguments, the first is the name of the file you want to open, test.txt. The second is the mode in which you want to open the file. The mode can be:

r = read only
w = read and write (the file is created, if it exists, it's truncated)
a = append (the file is created if it does not exist)

Now that we've opened the file, let's read the first line and display it's contents.















We used the readline() function to read from the file handle f.

We put the contents returned by readline() into the variable line.

Then we used the print() function to display the contents of the variable line.

Note the format for reading from the file into the variable.

line = f.readline()

Now let's read the next line and display it's contents.


















We call readline() again to read the next line.

We use the same variable line to put the contents of the next line.
tabs
Python automatically moved to the next line after we read the first one.

Now let's go to the beginning of the file and use a for loop to read all the lines and display them.


































To go to the start of the file, we use the seek() function. The argument to the seek() function is the position in the file that we want to go to. In our case, we use the argument 0. Which means, to the beginning of the file.

Notice that the file seems bigger than the one we have. This is because the print() function automatically adds a newline after it writes the line to the screen. But each line also has a newline character at the end. So the file looks like it's double-spaced.tabs

We can remove the newline character from the line before we print it to the screen like this.





















The line that removes the newline character from the file is:

line = line.strip()

What this does is remove leading and trailing whitespace. A newline character, or spaces, or tabs at the front or end of the line would be stripped.

Now let's close the file and do some writing exercises.





To open a file for writing, which will also create the file, do the following:




If the file writing-file.txt existed, then python will truncate it. So be careful not to use the name of a file that you already have. For example, if we had used the name test.txt, our file that we used for the reading exercise would have been deleted.

Now let's write a few lines into the new file.






We've written four lines into our new file. Now let's read them.










Wow! What happened? The error message indicates that the file was not opened for us to read. There's a way to open a file for both reading and writing, but the mode we used 'w' was for writing only.

Let's close the file and open it for both reading and writing.


Dictionaries

Python has a dictionary data type.

The dictionary data type stores items as KEY: VALUE pairs. Unlike a list that stores items by their position in the array:

a[0], a[1], a[2], a[3]... 

A dictionary has no concept of order. The programmer determines what the key is, and assigns a value to it.

Consider the following list:












We create an empty list:

l = list()

We then append elements to the list. Recall that the list is ordered from zero.

for x in range(10):
    l.append(x)

The range() function returns a list of integers. In our example, range(10) will return the list:
[0,1,2,3,4,5,6,7,8,9]

Now let's remove the third element (l[2] - since the list is ordered from zero).








Notice how the item at position 2 has gone.

Actually the pop(x) method removes the element which has the value x. In the example above, if there was no "2" in the list, then an error would have been returned. The way to remove the 2-nd element is to use the del method.








In the example above, the element with value "3" is in the index position [2].

The statement:

del l[2]

deletes the "3"

[0, 1, 3, 4, 5, 6, 7, 8, 9]

becomes

[0, 1, 4, 5, 6, 7, 8, 9]

Now let's insert the "3" and the "2" back.









The insert(Pos, Value) function takes two arguments. Pos = the position (zero-based) where we will insert the element, and Value = the value of the element that we're inserting.

And now for the "2"








The function call, l.insert(2,2) will insert the value = 2 as position = 2.

DICTIONARIES are different. They are not based on indexed positions. Each entry in a dictionary has a key that marks where it is. The value is the actual element.

We can create an empty dictionary the following way:







Just like the list() function creates an empty list, the dict() function creates an empty dictionary. We can also create an empty list the following way:

l = []

And we can create an empty dictionary the following way:

d = {}

A dictionary uses curly braces (also known as brackets) to indicate that it's a dictionary.

Now let's add some elements to our empty dictionary. Remember, we need to add a key and an element. Unlike a list that only requires that you specify the element that you want to insert or append. A dictionary has no concept of ordering. We will see later how to sort dictionaries.










Notice a couple of things:

  • The key can be anything. A string or a number.
  • The value can be anything. A string or a number.
  • The key does not have to be related to the value.
A dictionary has an update() method if you want to add more than one key:value pair at once.

Let's see how this works:










Note the statement:

d.update({4: 'four', 5: 'five', 'holiday': 'Kwanzaa', 'lunari': 'X'})

This adds four elements. The keys are:

  • 4
  • 5
  • holiday
  • lunari
The values of each of those keys are:
  • four
  • five
  • Kwanzaa
  • X
It's probably best practice to use the update() method for all inserts so that you get used to it for single or multiple inserts.

Since dictionaries are unordered, to get a list of all the keys, the dictionary provides a keys() method.

This is how it works:





And the dictionary object also has a values() method that retrieves all the values. It works like this:




The keys() method is typically used in a loop to print out all the values. Like this:












The loop iterates through the list. The list is a list of keys. And inside the loop, the following statement runs:

print(key, ': ', d[key])

The print() statement will print a comma-separated list of items and add a newline at the end.

Sunday, November 13, 2011

Lists

Python has a special type called a List. A Python list is a set of objects, they could be numbers, or strings, or other objects, enclosed in a pair of brackets, square brackets, [].


The Python range() function is a special function that returns a list object with a sequence of numbers. In our case, we asked for the range 1 to 5, excluding the last number, 5. If you don't specify the first number, the range will start at zero (0).


See that range(3,3) returns an empty list. That's because the end of the range isn't included and since we requested a range from 3 to 3, it's empty. The second range request (range(3,4)) returns a list of only one element, the number 3.

The range() function is commonly found in for loops.


We loop around five times, from 0, to 1, to 2, to 3, to 4. That's five times. In the first loop;

item = 0
value = (1 + 1) * 1 = 2

item = 1
value = (2 + 1) * 2 = 6

item = 2
value = (6 + 1) * 6 = 42

item = 3
value = (42 + 1) * 42 = 1806

item = 4
value = (1806 + 1) * 1806 = 3263442

One thing to remember about the range() function is that it can become inefficient with very large numbers. When used in a loop the way I've created it above, Python constructs  and allocates for a list object. A special type of list object known as a Tuple. We'll discuss this further later.

To find out how many items are in the list, you use the len() function, similar to a string.


Just as with the string object, you can iterate through the list using a for loop.


We added the numbers in the list using a for loop.

You can join two lists together using the "+", concatenation, operator. This is like the string concatenation operator. And just like the string repeat operator, "*", the same operator is used to make copies of lists.


The list "e" is the concatenation of lists "c" and "d". The list "f" is the list "c" repeated twice. This is similar to what we saw with strings.

And just like we saw how we can take slices of strings, we can also take slices of lists. For a comprehensive discussion of what the start and end indices mean, see the discusion on strings. For now, it's enough to say that the first element in the list has the index zero (0), and like strings, we can also count from the end of the list where the last element has the index -1.


a = [1, 2, 3, 4, 5, 6]

The number "1" has the index zero (0). The number "2" has the index 1. And so on. So the slice:

b = a[0:2]

Takes elements from index zero (0) -- the "1", up to and not including the index 2 -- the "3". So the slice becomes [1, 2].

But guess what, unlike strings, lists are mutable! You can change the list elements inline.


This wasn't possible with a string. In the case of strings, trying to change an element in the string by assigning it as:

s[index] = value

would give you an error message.

You can also delete and insert elements into the list. There are two ways to delete elements. The first is to assign an empty list in the position of the elements you want deleted. For example, say you have the list:

a = [1, 2, 3, 4, 5, 6]

and you want to delete the elements, [2, 3]. These are represented by slice a[1:3]. Here's how you do that.


Python also provides a delete function, called del, which makes this easier to see than empty list assignment.


Lastly let's discuss inserting elements into the list. To insert an element into a position in the list, you have to use slice notation. If you don't, unless you're adding elements at the end of the list, you will overwrite the value at the index.

For example, say you have the list:

a = [1, 2, 3, 4, 5, 6]

and you'd like to insert the value zero (0) right after the 3, you'd do the following:


Notice that slice [3:3] selects elements starting at position 3 (currently occupied by the number 4), but doesn't include that position. So, in effect what happens is that the element in that position is shifted over.

What if you wanted to replace the slice [2, 3, 4] with the single zero?


On the left side of the assignment, we selected the slice that includes [2, 3, 4], that's slice [1:4]. We replace that slice with the list [0].

One thing you have to be careful with about lists is copying them. Unlike strings, when you make assignments to lists, an alias is made.

Look at the following example.


Notice what happened. A list "a" was created with three elements, [1, 2, 3]. The variable "b" was set to "a". At this point, a copy of the list "a" wasn't made. "b" is a reference, or an alias, to the existing list created by "a".

So, when the element at index zero (0) is modified by the statement:

b[0] = 4

It's the same element that's referenced by "a". So when we print out the value of "a", we see that the first element, at index zero, has been changed.

How do you make copies?

Take an entire slice of the list.


In this example, we used the slice operator to make a copy of "a". Remember that when you omit the first parameter in the slice it's assumed that you're starting from the first element, and when you omit the last parameter, it's assumed that you're slicing all the way to the end. So, in this case, we're slicing from the first parameter, all the way to the end. Essentially a full copy of the list.

When we make assignments to the list "b" it doesn't affect the list "a" since we now have two separate lists.

Knowing when to slice and when to assign is very important, because when you pass a list item as an argument to a function, a reference to the list item is passed. If the function modifies the elements of the list, the modifications are global.

Here's an example.


In this example we have a list "a" with three elements, [1, 2, 3]. We then define a function delh() that takes a single argument. It's a list argument because in the function body we can see that it deletes the first element and then prints it.

In our example, the first call to the function is:

delh(a[:])

The argument a[:] is a copy of the list "a". The delh() function deletes the first element, and then prints the remaining elements.

Later, out of the function, when we print the values in "a" we see that it's unaffected. It's not changed.

In the second function call to delh() we do the following:

delh(a)

In this case, we're passing a reference to the list. The delh() function deletes the first element and prints the remaining ones. Later, out of the function, when we inspect our original list, "a", we find out that the first element was deleted.

A last word on lists. A couple of useful functions exists to convert strings to lists, and lists to strings.

The first is split. This is a string function that will split a string, based on whitespace, or any string, into a list.

Example:


In the first example, the list "l" is created using the following line:

l = string.split(s)

Because we haven't specified how to split the string, Python splits it based on whitespace. So we get a list of the words.

In the second split command,

ll = string.split(s, 's')

We instruct Python to split the string at each occurrence of the string, or letter, "s". The "s" won't be included in the result, but spaces will.

Now, if you have a list of elements, you can create a string using the "join" function.


In the first join statement, since we haven't instructed Python how we'd want the list join performed, Python joins the elements of the list using a single space. However, in the second example, we ask Python to join the elements using the string "::", two colons.

This has been a fairly long section on lists. The next one is very short, because most of the material covering lists also pertains to that special type of a list, called a Tuple.

String operations - Part 3 of 3

Python Strings are immutable. Which means that once you create one, you can't change it. In part 2 of our discussion on strings, we saw how to extract a character from a string.


However, it's illegal to try and replace a character inline. The string "a" in the example above cannot be changed.


You get the error message displayed above. That the 'str" object does not support assignment. If you want to change the contents of the string, the only thing to do is create another one. Using slices, you can do the following:


What we did is take the letter "m", add, or concatenate, it to the rest of the original string "a" starting at position 1 ("ack") and then assign it to a new string called "a." Looks like we modified the original string "a" but in effect what Python does is destroy the old string "a" and create a new one. This allows us to do the following:


In each of the assignments above, the original string "a" is destroyed, a new string is created with the characters on the right side, and that's assigned to the variable "a." So, a="jack" is a different string from a="jill".

We've already seen how to find the length of a string using the Python function len().


We can also use the for loop in Python to count the characters.


The string module has a find() function. This function allows you to find a character, or another string, inside a string. In order to use the find() function, we have to import the string module.


Notice how we call the find() function using the string module class identifier. The find() function is called a class function. It requires two arguments; the string that you're searching and the substring to look for.

In our example:

string.find(a, "n")

This instructs the find() function to look for the "n" string inside the "a" string. The "a" variable points to the string "canada". In that string, "n" is at position 2, remembering that the first character is at position 0.

If the find() function does not find the string that we're looking for, it returns the number -1.


See that the value of "i", in the second find() is -1 because the string "canada" does not contain the letter "q" that we're looking for.

The find() function is not limited to looking for simple characters, it can also find entire words.


In the find() function call above, the word "canada" is at position 36 in the string "a". The number 36 is the starting position of the word "canada". It's the position where the "c" in "canada" is.


In the example above, after finding the position of the string "canada" we extracted it. Not a very useful example but illustrative of the fact that you can use an example of this sort to pull out words, based on whitespace.

The string module contains some useful functions that will help when analysing text.


How would you use these? Well, you can use the string.lowercase variable to test if a letter in your particular string is lowercase. Here's an example that counts lowercase and uppercase characters in a string.


We have our string with upper and lowercase characters, "s". We then use a for loop to loop through each character. If its a lowercase character, we increment the value of the lowercount variable. If its an uppercase character, we increment the value of the uppercount variable. Notice how we use the find() function.

string.find(string.lowercase, item)

"item" is the character we're looking for.
string.lowercase is the string we're searching in. string.lowercase contains all the lowercase characters. If we can't find the character, the find() function returns the value -1. Otherwise it returns a value from zero (0) to the length of the string.

There is another, more elegant, way to do something like this. That's using the string "in" operator. The "in" operator checks to see if one string is "in"side another. For example, if the string is "canada" and we issue the following statement:

a = "can" in "canada"

The value of "a" will be True. Because the string "can" is inside "canada"

Here's the counting program, written using the "in" operator.



The string module has a number of interesting functions that you can use to manipulate strings. Functions to modify strings from uppercase to lowercase and vice versa. To split strings at whitespace into words. To convert strings to numbers.