Text statistics
Type several lines of text and see its counts (lines, words, letters, average and longest word), its most common words, a chart of its letters, the palindromes in it - words and whole lines - and the groups of words that are anagrams of each other, from a text menu.
Every screen below was recorded under CPython. When this page was built, the EML interpreter replayed each session from the same input and printed the same bytes.
About
Type several lines of text - a line with just a dot ends it - and look at it five ways: its counts (lines, words, letters, characters, the average and the longest word), its most common words, a chart of its letters, the palindromes in it, and the groups of words in it that are anagrams of each other. New text replaces the old; up to 200 lines are kept.
main.eml- the menu, reading the text, and what the screen showsanalyze.eml- splitting lines into words, counting words and letters, sorting by count, and finding palindromes and anagramstext.eml- trimming, lower case, letters and digits, reversing, sorting the letters of a word, and rounding
A word is a run of letters, digits and apostrophes, with apostrophes at either end taken off - "don't" is one word, 'tis is "tis" - and words are compared in lower case. The common words are sorted by count, largest first, and words with the same count in alphabetical order, so the list is the same every time. Each letter of the chart has its count, its share of all letters to one decimal (rounded half up in whole numbers) and a bar scaled so the most common letter is 40 wide.
A palindrome word has at least three letters and reads the same backwards (bob, kayak, noon, level); a palindrome line is one whose letters, ignoring case, spaces and punctuation, read the same backwards ("Was it a car or a cat I saw?"). An anagram group is two or more different words of at least three letters made of exactly the same letters (listen, silent, enlist): sorting the letters of each word gives the same key. The interpreter that checks every session does not run string or list methods yet, so reversing, sorting and joining are written out by hand.
What is checked: a menu choice is one of 1 to 7, and anything but entering text needs some text first; an empty text (just the dot) leaves the old text as it was.
Sessions: sessions/basic.in types five lines with three anagram groups, four palindrome words and two palindrome lines, then shows every view; sessions/bad-input.in asks for views before any text, picks numbers that are not on the menu, ends a text before it starts, then types numbers, a line of dots, apostrophes and quotes, and an empty line - a text with no palindromes and no anagrams.
Built on the verified corpus cases word-frequency-counter (counting words in order of first appearance), char-frequency-table (counting characters, with a bar for each), palindrome-checker (a word against its reverse) and anagram-checker (two words made of the same letters).
Recorded sessions
What the screen shows while someone uses the program. Each typed line appears after its prompt, the way a terminal shows it.
bad-input
interpreter: byte-equal
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 2
Enter some text first.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 6
Enter some text first.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 8
Pick a number from 1 to 7.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 23
Pick a number from 1 to 7.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 1
Type the text, one line at a time; a line with just . ends it.
> .
No text entered; the text is as it was.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 1
Type the text, one line at a time; a line with just . ends it.
> 123 456
> ...
> 'tis the "end" - don't stop
>
> .
Read 4 lines, 7 words.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 2
-- counts --
lines 4 (3 with text)
words 7
letters 17
characters 39 (with spaces and punctuation)
average word 3.4 characters
longest word don't (5 characters)
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 3
-- common words --
7 words, 7 different; the most common:
1 123
1 456
1 don't
1 end
1 stop
1 the
1 tis
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 4
-- letters (17 in all) --
d 2 11.8% ####################
e 2 11.8% ####################
h 1 5.9% ##########
i 1 5.9% ##########
n 2 11.8% ####################
o 2 11.8% ####################
p 1 5.9% ##########
s 2 11.8% ####################
t 4 23.5% ########################################
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 5
-- palindromes --
words: none
lines: none
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 6
-- anagrams --
none
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 7
Bye.
What was typed (18 lines)
2
6
8
23
1
.
1
123 456
...
'tis the "end" - don't stop
.
2
3
4
5
6
7
basic
interpreter: byte-equal
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 1
Type the text, one line at a time; a line with just . ends it.
> Listen to the silent night; enlist the stars.
> Was it a car or a cat I saw?
> Evil and vile are the same letters as live.
> Bob saw a kayak at noon, and the level of the river rose.
> Never odd or even.
> .
Read 5 lines, 43 words.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 2
-- counts --
lines 5 (5 with text)
words 43
letters 146
characters 191 (with spaces and punctuation)
average word 3.4 characters
longest word letters (7 characters)
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 3
-- common words --
43 words, 34 different; the most common:
5 the
3 a
2 and
2 or
2 saw
1 are
1 as
1 at
1 bob
1 car
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 4
-- letters (146 in all) --
a 17 11.6% ##############################
b 2 1.4% ###
c 2 1.4% ###
d 4 2.7% #######
e 23 15.8% ########################################
f 1 0.7% ##
g 1 0.7% ##
h 6 4.1% ##########
i 10 6.8% #################
k 2 1.4% ###
l 9 6.2% ################
m 1 0.7% ##
n 10 6.8% #################
o 9 6.2% ################
r 10 6.8% #################
s 12 8.2% #####################
t 16 11.0% ############################
v 7 4.8% ############
w 3 2.1% #####
y 1 0.7% ##
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 5
-- palindromes --
words: bob, kayak, noon, level
lines:
Was it a car or a cat I saw?
Never odd or even.
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 6
-- anagrams --
listen, silent, enlist
was, saw
evil, vile, live
== Text statistics ==
1) enter text 2) counts 3) common words 4) letter chart
5) palindromes 6) anagrams 7) quit
choice> 7
Bye.
What was typed (13 lines)
1
Listen to the silent night; enlist the stars.
Was it a car or a cat I saw?
Evil and vile are the same letters as live.
Bob saw a kayak at noon, and the level of the river rose.
Never odd or even.
.
2
3
4
5
6
7
Modules
The program as written, entry module first. Each module transpiles to its own Python file, which is what eml project run executes.
main.eml(entry)
eml# P016 text statistics: type several lines of text, then look at its counts,
# its most common words, a chart of its letters, its palindromes and its
# anagrams. New text replaces the old.
import analyze
import text
200 => most_lines
10 => most_common
40 => widest_bar
def plural(n, word):
if n == 1:
return "1 " + word
return str(n) + " " + word + "s"
def row(label, value):
return " " + ("%-16s" % label) + ("%6s" % value)
def joined(items):
"" => out
for i in [0:len(items) - 1]:
if i > 0:
out + ", " => out
out + items[i] => out
return out
def read_text():
# Lines typed until a line that is just a dot; at most most_lines kept.
"Type the text, one line at a time; a line with just . ends it." ^0
[] => lines
False => full
while True:
input("> ") => line
if text.trim(line) == ".":
return lines
if len(lines) < most_lines:
lines + [line] => lines
elif not full:
("At most " + str(most_lines) + " lines; the rest is left out.") ^0
True => full
def show_counts(lines):
analyze.all_words(lines) => ws
0 => with_text
0 => chars
0 => letters
for line in lines:
chars + len(line) => chars
letters + len(analyze.letters_only(line)) => letters
if text.trim(line) != "":
with_text + 1 => with_text
0 => word_chars
"" => longest
for w in ws:
word_chars + len(w) => word_chars
if len(w) > len(longest):
w => longest
"" ^0
"-- counts --" ^0
(row("lines", str(len(lines))) + " (" + str(with_text) + " with text)") ^0
row("words", str(len(ws))) ^0
row("letters", str(letters)) ^0
(row("characters", str(chars)) + " (with spaces and punctuation)") ^0
if len(ws) == 0:
row("average word", "-") ^0
row("longest word", "-") ^0
else:
(row("average word", text.tenths(word_chars, len(ws))) + " characters") ^0
(" " + ("%-16s" % "longest word") + longest + " (" + str(len(longest)) + " characters)") ^0
def show_common(lines):
analyze.counted(analyze.all_words(lines)) => pairs
"" ^0
"-- common words --" ^0
if len(pairs) == 0:
" (no words)" ^0
else:
0 => total
for p in pairs:
total + p[1] => total
(" " + str(total) + " words, " + str(len(pairs)) + " different; the most common:") ^0
0 => shown
for p in analyze.by_count(pairs):
if shown < most_common:
(("%6d" % p[1]) + " " + p[0]) ^0
shown + 1 => shown
def show_letters(lines):
analyze.letter_counts(lines) => counts
0 => total
0 => most
for p in counts:
total + p[1] => total
if p[1] > most:
p[1] => most
"" ^0
if total == 0:
"-- letters --" ^0
" (no letters)" ^0
else:
("-- letters (" + str(total) + " in all) --") ^0
for p in counts:
(" " + p[0] + ("%5d" % p[1]) + ("%7s" % (text.tenths(100 * p[1], total) + "%")) + " " + ("#" * text.half_up(widest_bar * p[1], most))) ^0
def show_palindromes(lines):
analyze.palindrome_words(analyze.all_words(lines)) => ws
analyze.palindrome_lines(lines) => ls
"" ^0
"-- palindromes --" ^0
if len(ws) == 0:
" words: none" ^0
else:
(" words: " + joined(ws)) ^0
if len(ls) == 0:
" lines: none" ^0
else:
" lines:" ^0
for line in ls:
(" " + line) ^0
def show_anagrams(lines):
analyze.anagram_groups(analyze.all_words(lines)) => groups
"" ^0
"-- anagrams --" ^0
if len(groups) == 0:
" none" ^0
for g in groups:
(" " + joined(g)) ^0
[] => lines
True => running
while running:
"" ^0
"== Text statistics ==" ^0
"1) enter text 2) counts 3) common words 4) letter chart" ^0
"5) palindromes 6) anagrams 7) quit" ^0
text.trim(input("choice> ")) => choice
if choice == "1":
read_text() => typed
if len(typed) == 0:
"No text entered; the text is as it was." ^0
else:
typed => lines
("Read " + plural(len(lines), "line") + ", " + plural(len(analyze.all_words(lines)), "word") + ".") ^0
elif choice == "7":
False => running
elif len(choice) == 1 and choice in "23456" and len(lines) == 0:
"Enter some text first." ^0
elif choice == "2":
show_counts(lines)
elif choice == "3":
show_common(lines)
elif choice == "4":
show_letters(lines)
elif choice == "5":
show_palindromes(lines)
elif choice == "6":
show_anagrams(lines)
else:
"Pick a number from 1 to 7." ^0
"Bye." ^0
Python projection (main.py)
import analyze
import text
most_lines = 200
most_common = 10
widest_bar = 40
def plural(n, word):
if n == 1:
return "1 " + word
return str(n) + " " + word + "s"
def row(label, value):
return " " + "%-16s" % label + "%6s" % value
def joined(items):
out = ""
for i in range(0, len(items)):
if i > 0:
out = out + ", "
out = out + items[i]
return out
def read_text():
print("Type the text, one line at a time; a line with just . ends it.")
lines = []
full = False
while True:
line = input("> ")
if text.trim(line) == ".":
return lines
if len(lines) < most_lines:
lines = lines + [line]
elif not full:
print("At most " + str(most_lines) + " lines; the rest is left out.")
full = True
def show_counts(lines):
ws = analyze.all_words(lines)
with_text = 0
chars = 0
letters = 0
for line in lines:
chars = chars + len(line)
letters = letters + len(analyze.letters_only(line))
if text.trim(line) != "":
with_text = with_text + 1
word_chars = 0
longest = ""
for w in ws:
word_chars = word_chars + len(w)
if len(w) > len(longest):
longest = w
print("")
print("-- counts --")
print(row("lines", str(len(lines))) + " (" + str(with_text) + " with text)")
print(row("words", str(len(ws))))
print(row("letters", str(letters)))
print(row("characters", str(chars)) + " (with spaces and punctuation)")
if len(ws) == 0:
print(row("average word", "-"))
print(row("longest word", "-"))
else:
print(row("average word", text.tenths(word_chars, len(ws))) + " characters")
print(" " + "%-16s" % "longest word" + longest + " (" + str(len(longest)) + " characters)")
def show_common(lines):
pairs = analyze.counted(analyze.all_words(lines))
print("")
print("-- common words --")
if len(pairs) == 0:
print(" (no words)")
else:
total = 0
for p in pairs:
total = total + p[1]
print(" " + str(total) + " words, " + str(len(pairs)) + " different; the most common:")
shown = 0
for p in analyze.by_count(pairs):
if shown < most_common:
print("%6d" % p[1] + " " + p[0])
shown = shown + 1
def show_letters(lines):
counts = analyze.letter_counts(lines)
total = 0
most = 0
for p in counts:
total = total + p[1]
if p[1] > most:
most = p[1]
print("")
if total == 0:
print("-- letters --")
print(" (no letters)")
else:
print("-- letters (" + str(total) + " in all) --")
for p in counts:
print(" " + p[0] + "%5d" % p[1] + "%7s" % (text.tenths(100 * p[1], total) + "%") + " " + "#" * text.half_up(widest_bar * p[1], most))
def show_palindromes(lines):
ws = analyze.palindrome_words(analyze.all_words(lines))
ls = analyze.palindrome_lines(lines)
print("")
print("-- palindromes --")
if len(ws) == 0:
print(" words: none")
else:
print(" words: " + joined(ws))
if len(ls) == 0:
print(" lines: none")
else:
print(" lines:")
for line in ls:
print(" " + line)
def show_anagrams(lines):
groups = analyze.anagram_groups(analyze.all_words(lines))
print("")
print("-- anagrams --")
if len(groups) == 0:
print(" none")
for g in groups:
print(" " + joined(g))
lines = []
running = True
while running:
print("")
print("== Text statistics ==")
print("1) enter text 2) counts 3) common words 4) letter chart")
print("5) palindromes 6) anagrams 7) quit")
choice = text.trim(input("choice> "))
if choice == "1":
typed = read_text()
if len(typed) == 0:
print("No text entered; the text is as it was.")
else:
lines = typed
print("Read " + plural(len(lines), "line") + ", " + plural(len(analyze.all_words(lines)), "word") + ".")
elif choice == "7":
running = False
elif len(choice) == 1 and choice in "23456" and len(lines) == 0:
print("Enter some text first.")
elif choice == "2":
show_counts(lines)
elif choice == "3":
show_common(lines)
elif choice == "4":
show_letters(lines)
elif choice == "5":
show_palindromes(lines)
elif choice == "6":
show_anagrams(lines)
else:
print("Pick a number from 1 to 7.")
print("Bye.")
analyze.eml
eml# P016 text statistics - what is counted and found in the text, which is a
# list of lines. A word is a run of letters, digits and apostrophes, with
# apostrophes at either end taken off ("don't" is one word, "'tis'" is
# "tis"), and words are compared in lower case.
import text
def words_of(line):
[] => out
"" => word
for c in line + " ":
if text.is_letter(c) or text.is_digit(c) or c == "'":
word + c => word
else:
while len(word) > 0 and word[0] == "'":
word[1:len(word)] => word
while len(word) > 0 and word[len(word) - 1] == "'":
word[0:len(word) - 1] => word
if word != "":
out + [text.lower(word)] => out
"" => word
return out
def all_words(lines):
[] => out
for line in lines:
out + words_of(line) => out
return out
def letters_only(s):
# The letters of s in lower case, everything else left out.
"" => out
for c in s:
if text.is_letter(c):
out + text.lower_char(c) => out
return out
def counted(words):
# [word, count] for every distinct word, in order of first appearance.
{} => seen
[] => out
for w in words:
if w in seen:
out[seen[w]][1] + 1 => out[seen[w]][1]
else:
len(out) => seen[w]
out + [[w, 1]] => out
return out
def by_count(pairs):
# pairs sorted by count, largest first, equal counts in alphabetical order
# (insertion sort, so the order of ties is fixed).
[] => out
for p in pairs:
len(out) => i
out + [p] => out
while i > 0 and (out[i - 1][1] < p[1] or (out[i - 1][1] == p[1] and out[i - 1][0] > p[0])):
out[i - 1] => out[i]
i - 1 => i
p => out[i]
return out
def letter_counts(lines):
# [letter, count] for each of a-z that occurs, in alphabetical order.
[] => counts
for i in [0:25]:
counts + [[text.lower_letters[i], 0]] => counts
for line in lines:
for c in letters_only(line):
for i in [0:25]:
if text.lower_letters[i] == c:
counts[i][1] + 1 => counts[i][1]
[] => out
for p in counts:
if p[1] > 0:
out + [p] => out
return out
def is_palindrome(s):
return s == text.reverse(s)
def palindrome_words(words):
# Distinct words of three or more letters that read the same backwards.
[] => out
for p in counted(words):
if len(p[0]) >= 3 and letters_only(p[0]) == p[0] and is_palindrome(p[0]):
out + [p[0]] => out
return out
def palindrome_lines(lines):
# Lines whose letters, ignoring case and everything else, read the same
# backwards - at least three letters of them.
[] => out
for line in lines:
letters_only(line) => s
if len(s) >= 3 and is_palindrome(s):
out + [text.trim(line)] => out
return out
def anagram_groups(words):
# Groups of two or more distinct words of three or more letters made of
# the same letters, in order of first appearance.
{} => where
[] => groups
for p in counted(words):
p[0] => w
if len(w) >= 3 and letters_only(w) == w:
text.sorted_chars(w) => key
if key in where:
groups[where[key]] + [w] => groups[where[key]]
else:
len(groups) => where[key]
groups + [[w]] => groups
[] => out
for g in groups:
if len(g) >= 2:
out + [g] => out
return out
Python projection (analyze.py)
import text
def words_of(line):
out = []
word = ""
for c in line + " ":
if text.is_letter(c) or text.is_digit(c) or c == "'":
word = word + c
else:
while len(word) > 0 and word[0] == "'":
word = word[1:len(word)]
while len(word) > 0 and word[len(word) - 1] == "'":
word = word[0:len(word) - 1]
if word != "":
out = out + [text.lower(word)]
word = ""
return out
def all_words(lines):
out = []
for line in lines:
out = out + words_of(line)
return out
def letters_only(s):
out = ""
for c in s:
if text.is_letter(c):
out = out + text.lower_char(c)
return out
def counted(words):
seen = {}
out = []
for w in words:
if w in seen:
out[seen[w]][1] = out[seen[w]][1] + 1
else:
seen[w] = len(out)
out = out + [[w, 1]]
return out
def by_count(pairs):
out = []
for p in pairs:
i = len(out)
out = out + [p]
while i > 0 and (out[i - 1][1] < p[1] or out[i - 1][1] == p[1] and out[i - 1][0] > p[0]):
out[i] = out[i - 1]
i = i - 1
out[i] = p
return out
def letter_counts(lines):
counts = []
for i in range(0, 26):
counts = counts + [[text.lower_letters[i], 0]]
for line in lines:
for c in letters_only(line):
for i in range(0, 26):
if text.lower_letters[i] == c:
counts[i][1] = counts[i][1] + 1
out = []
for p in counts:
if p[1] > 0:
out = out + [p]
return out
def is_palindrome(s):
return s == text.reverse(s)
def palindrome_words(words):
out = []
for p in counted(words):
if len(p[0]) >= 3 and letters_only(p[0]) == p[0] and is_palindrome(p[0]):
out = out + [p[0]]
return out
def palindrome_lines(lines):
out = []
for line in lines:
s = letters_only(line)
if len(s) >= 3 and is_palindrome(s):
out = out + [text.trim(line)]
return out
def anagram_groups(words):
where = {}
groups = []
for p in counted(words):
w = p[0]
if len(w) >= 3 and letters_only(w) == w:
key = text.sorted_chars(w)
if key in where:
groups[where[key]] = groups[where[key]] + [w]
else:
where[key] = len(groups)
groups = groups + [[w]]
out = []
for g in groups:
if len(g) >= 2:
out = out + [g]
return out
text.eml
eml# P016 text statistics - characters and strings. The interpreter that checks
# every session does not run string or list methods yet, so lower case,
# reversing, sorting and joining are written out here.
"abcdefghijklmnopqrstuvwxyz" => lower_letters
"ABCDEFGHIJKLMNOPQRSTUVWXYZ" => upper_letters
def trim(s):
# s without the spaces at either end.
0 => i
len(s) => j
while i < j and s[i] == " ":
i + 1 => i
while j > i and s[j - 1] == " ":
j - 1 => j
return s[i:j]
def lower_char(c):
for i in [0:25]:
if upper_letters[i] == c:
return lower_letters[i]
return c
def lower(s):
"" => out
for c in s:
out + lower_char(c) => out
return out
def is_letter(c):
return c in lower_letters or c in upper_letters
def is_digit(c):
return c in "0123456789"
def reverse(s):
"" => out
for c in s:
c + out => out
return out
def sorted_chars(s):
# The characters of s in order, as a string (insertion sort).
[] => cs
for c in s:
len(cs) => i
cs + [c] => cs
while i > 0 and cs[i - 1] > c:
cs[i - 1] => cs[i]
i - 1 => i
c => cs[i]
"" => out
for c in cs:
out + c => out
return out
def half_up(n, d):
# n / d rounded half up to a whole number, for n >= 0 and d > 0.
2 * n + d => t
return int((t - t % (2 * d)) / (2 * d))
def tenths(n, d):
# n / d to one decimal, rounded half up, as text: 19 / 5 is "3.8".
half_up(10 * n, d) => t
return str(int((t - t % 10) / 10)) + "." + str(t % 10)
Python projection (text.py)
lower_letters = "abcdefghijklmnopqrstuvwxyz"
upper_letters = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"
def trim(s):
i = 0
j = len(s)
while i < j and s[i] == " ":
i = i + 1
while j > i and s[j - 1] == " ":
j = j - 1
return s[i:j]
def lower_char(c):
for i in range(0, 26):
if upper_letters[i] == c:
return lower_letters[i]
return c
def lower(s):
out = ""
for c in s:
out = out + lower_char(c)
return out
def is_letter(c):
return c in lower_letters or c in upper_letters
def is_digit(c):
return c in "0123456789"
def reverse(s):
out = ""
for c in s:
out = c + out
return out
def sorted_chars(s):
cs = []
for c in s:
i = len(cs)
cs = cs + [c]
while i > 0 and cs[i - 1] > c:
cs[i] = cs[i - 1]
i = i - 1
cs[i] = c
out = ""
for c in cs:
out = out + c
return out
def half_up(n, d):
t = 2 * n + d
return int((t - t % (2 * d)) / (2 * d))
def tenths(n, d):
t = half_up(10 * n, d)
return str(int((t - t % 10) / 10)) + "." + str(t % 10)