Roget's Thesaurus — A Closer Reading
Edition facts
For Roget's Thesaurus — A Closer Reading, the stored edition analysis reports 206,199 words, 14 hr 57 min estimated reading time, and 26 detected text sections.
The text analysis averages about 12.8 words per sentence, while the detected sections provide another way to judge how the source is divided.
Project Gutenberg metadata also associates the work with “English language -- Synonyms and antonyms,” connecting these edition facts with the source record’s subject description.
Calculated from edition completeness, EPUB availability, text structure and catalogue metadata. Not a user rating.
How this score is calculated
- Description quality20 pts
- Title & short description10 pts
- Source metadata20 pts
- Text length15 pts
- Chapters / structure15 pts
- EPUB file integrity20 pts
Total of 100 points, scaled to a 2.5-5.0 range. Editions with an empty description or a missing EPUB file are not scored.
Read the Text
English words and phrases classified and arranged so as to facilitate the expression of ideas and assist in literary composition
An electronic thesaurus derived from the version of Roget's Thesaurus published in 1911.
MICRA, Inc. makes no representation that the original 1911 printed work on which this is based is now in the public domain in any particular country. However, MICRA, Inc. makes no proprietary claims regarding this electronic version of the 1911 thesaurus. If the 1911 work is currently public domain, this electronic version can also be treated as public domain.
Note that this version of Thesaurus-1911 has been supplemented with over 1,000 words not present in the original 1911 edition, but many modern words are still missing. About 1500 verbs (out of 6500) which can be found in an 80,000-word spell-checker are absent from this work. The deficiency of nouns is probably much worse, especially on technical topics. Of 40,000 unique words contained in the original text, 12,000 are not recognized by a spell-checker. Most of these are foreign words (primarily Latin), and many are obsolete. In this version, these words are marked as such by comments in square brackets. Although this version has been proof-read, there are doubtless numerous residual transcription errors, some of which may be obvious even without reference to the original text. We will be grateful if any of these are brought to our attention; the corrections will appear in subsequent versions.
The original arrangement has also been modified slightly in several places, in particular by splitting one entry into two. A version of the 1911 thesaurus which is almost identical to the original (only a small number of additions to the original work) has also been prepared by MICRA, Inc., and also carries no restrictions from MICRA. Copies of that version or this one may be purchased for $40.00 from MICRA, Inc., or from the Austin Code Works, Austin Texas.
Occasional references to numbers starting with "@" are the embryonic beginnings of a reorganized version, mentioned below. A few comments are also included within curly brackets {}.
The following additional differences will be noted between this version and the original edition of the printed 1911 thesaurus:
(1) the space-saving abbreviations in the original, using hyphens to represent common words, prefixes or suffixes, have been expanded into the full words or phrases.
(2) the side-by-side format for words and their opposites has been abandoned. Words are listed in order of their entry number.
(3) each main entry (1035 entries) has a pound sign "#" in front of the number to facilitate computerized search.
(4) where italics occurred in the original, italics are used in the Microsoft Word format file. In the plain ASCII file, this formatting is lost.
(5) in the original book, words which were obsolete (in 1911) were marked with a dagger. In this version, those words are marked with a vertical bar ("|").
Some of the words which were still current in 1911, but are no longer found in a current college-size dictionary (presently obsolete words), or which are no longer used in the specific indicated sense, have been marked with a bar followed by an exclamation point "|!". However, this marking process has just commenced, and only a small portion of the words which are now obsolete have been thus marked. Most though not all of the foreign-language phrases are now obsolete. The "obsolete" notation [obs3] indicates that the previous word (or some word in the previous phrase) is not recognized by the word processor's spelling checker, and also is either NOT in a modern college-sized dictionary, or is noted there as being "ARCHAIC".
(7) This file contains only the main body of the thesaurus. Neither outline nor index are contained here. The outline with an overview of the organization of the concepts is contained in a separate file, "outline.doc", on the distribution disk.
This first edition of this supplemented 1911 thesaurus (June 1991) is very much less complete than the latest editions of commercial thesauri, and is probably not suitable for use as an adjunct to word-processing programs, but it has no proprietary claims attached to it by MICRA, Inc., and does not contain any material published commercially after 1911.
The 1911 edition of Roget's Thesaurus, as presented in this digital version, is a work of systematic classification. Its entries are numbered and cross-referenced, with each category introduced by a bold headword. The text alternates between dense clusters of synonyms—often separated by semicolons—and occasional explanatory notes in square brackets. This structure creates a distinctive rhythm: rapid-fire lists of words punctuated by brief editorial interventions.
Taxonomic Architecture and the Pace of Lists
The thesaurus is organized into numbered sections, each covering a conceptual domain. For example, section #558 on engraving lists dozens of terms—'line engraving, mezzotint engraving, stipple engraving, chalk engraving'—in a single paragraph. The pace is relentless, with no full stops until the end of the entry. This accumulation mimics the associative nature of language, but it also demands careful reading. The editor has expanded abbreviations from the original print edition, which slightly slows the rhythm but improves clarity. The use of semicolons to separate subgroups within a list creates a subtle hierarchy, guiding the reader through related terms without breaking the flow.
Annotations and the Voice of the Editor
Square brackets mark editorial comments, such as '[obs3]' for obsolete words or '[Fr]' for foreign terms. These annotations introduce a second voice into the text—one that is scholarly and precise. For instance, the entry for 'sculpture' includes 'insculpture| [obs3]', indicating a word marked with a vertical bar in the original. The editor also notes that 'about 1500 verbs ... are absent from this work' and that '12,000 [words] are not recognized by a spell-checker.' These asides break the monotony of the lists and provide context for the reader. They also reveal the editor's awareness of the thesaurus's limitations, acknowledging that 'many modern words are still missing.'
Shifts in Density: From Technical Terms to Philosophical Quotations
While most entries are dense with synonyms, some sections shift to a more discursive tone. The entry on 'Language' (#560) includes a list of related fields—'lexicology, philology, glossology, glottology'—but also a quotation from Selden: 'syllables govern the world.' This sudden inclusion of a philosophical remark alters the pace, offering a moment of reflection amid the catalog. Similarly, the section on 'Artist' (#559) ends with the phrase 'photo safari; "with gun and camera"', which feels almost anecdotal. These variations prevent the text from becoming purely mechanical, hinting at the human curiosity behind the classification.
Readers approaching this digital edition should note that the original 1911 arrangement has been modified slightly, with one entry split into two. The editor has also added over 1,000 words not present in the original. To navigate effectively, pay attention to the pound signs (#) marking main entries and the square brackets that signal editorial commentary. The thesaurus rewards both targeted searches and casual browsing, revealing the unexpected connections that Roget's system makes visible.
I finished that Roget volume and kept turning back to its strange, quiet order—how the editor gently sorted obsolete words like pressed flowers. It reminded me of another patient archive, The Gutenberg Webster's Unabridged Dictionary: Section F, G and H — Background and Themes, where definitions sprawl and breathe. There’s a tenderness in old reference books, a feeling that someone wanted every word to find a home.
There are no reviews for this eBook.
Record your reading impressions
A brief reflection can help important ideas stay with you longer.