Перейти к содержимому

Result search python что это

  • автор:

re introduction

This chapter gives an introduction to the re module. This module is part of the standard library. For some examples, the equivalent normal string method is also shown for comparison. This chapter just focuses on the basics of using functions from the re module. Regular expression features will be covered from the next chapter onwards.

re module documentation

It is always a good idea to know where to find the documentation. The default offering for Python regular expressions is the re standard library module. Visit docs.python: re for information on available methods, syntax, features, examples and more. Here’s a quote:

A regular expression (or RE) specifies a set of strings that matches it; the functions in this module let you check if a particular string matches a given regular expression

re.search()

Normally you’d use the in operator to test whether a string is part of another string or not. For regular expressions, use the re.search() function whose argument list is shown below.

The first argument is the RE pattern you want to test against the input string, which is the second argument. flags is optional, it helps to change the default behavior of RE patterns.

As a good practice, always use raw strings to construct the RE pattern. This will become clearer in later chapters. Here are some examples to get started.

Before using the re module, you need to import it. Further example snippets will assume that this module is already loaded. The return value of the re.search() function is a re.Match object when a match is found and None otherwise (note that I treat re as a word, not as r and e separately, hence the use of a instead of an). More details about the re.Match object will be discussed in the Working with matched portions chapter. For presentation purposes, the examples will use the bool() function to show True or False depending on whether the RE pattern matched or not.

Here’s an example with the flags optional argument. By default, the pattern will match the input string case sensitively. By using the re.I flag, you can match case insensitively. See the Flags chapter for more details.

re.search() in conditional expressions

As Python evaluates None as False in boolean context, re.search() can be used directly in conditional expressions. See also docs.python: Truth Value Testing.

Here are some examples with list comprehensions and generator expressions:

re.sub()

For normal search and replace, you’d use the str.replace() method. For regular expressions, use the re.sub() function, whose argument list is shown below.

re.sub(pattern, repl, string, count=0, flags=0)

The first argument is the RE pattern to match against the input string, which is the third argument. The second argument specifies the string which will replace the portions matched by the RE pattern. count and flags are optional arguments.

A common mistake, not specific to re.sub() , is forgetting that strings are immutable in Python.

Compiling regular expressions

Regular expressions can be compiled using the re.compile() function, which gives back a re.Pattern object.

The top level re module functions are all available as methods for such objects. Compiling a regular expression is useful if the RE has to be used in multiple places or called upon multiple times inside a loop (speed benefit).

By default, Python maintains a small list of recently used RE, so the speed benefit doesn’t apply for trivial use cases. See also stackoverflow: Is it worth using re.compile?

Some of the methods available for compiled patterns also accept more arguments than those available for the top level functions of the re module. For example, the search() method on a compiled pattern has two optional arguments to specify the start and end index positions. Similar to the range() function and slicing notation, the ending index has to be specified 1 greater than the desired index.

Note that there’s no flags option as that has to be specified with re.compile() .

bytes

To work with the bytes data type, the RE must be specified as bytes as well. Similar to the str RE, use raw format to construct a bytes RE.

re(gex)? playground

To make it easier to experiment, I wrote an interactive TUI app. See PyRegexPlayground repo for installation instructions and usage guide. A sample screenshot is shown below:

Cheatsheet and Summary

This chapter introduced the re module, which is part of the standard library. Functions re.search() and re.sub() were discussed as well as how to compile RE using the re.compile() function. The RE pattern is usually defined using raw strings. For byte input, the pattern has to be of byte type too. Although the re module is good enough for most use cases, there are situations where you need to use the third-party regex module. To avoid mixing up features, a separate chapter is dedicated for the regex module at the end of this book.

The next section has exercises to test your understanding of the concepts introduced in this chapter. Please do solve them before moving on to the next chapter.

Exercises

Try to solve exercises in every chapter using only the features discussed until that chapter. Some of the exercises will be easier to solve with techniques presented in later chapters, but the aim of these exercises is to explore the features presented so far.

All the exercises are also collated together in one place at Exercises.md. For solutions, see Exercise_solutions.md.

a) Check whether the given strings contain 0xB0 . Display a boolean result as shown below.

b) Replace all occurrences of 5 with five for the given string.

c) Replace only the first occurrence of 5 with five for the given string.

d) For the given list, filter all elements that do not contain e .

e) Replace all occurrences of note irrespective of case with X .

f) Check if at is present in the given byte input data.

g) For the given input string, display all lines not containing start irrespective of case.

h) For the given list, filter all elements that contain either a or w .

i) For the given list, filter all elements that contain both e and n .

j) For the given string, replace 0xA0 with 0x7F and 0xC0 with 0x1F .

re — Regular expression operations¶

This module provides regular expression matching operations similar to those found in Perl.

Both patterns and strings to be searched can be Unicode strings ( str ) as well as 8-bit strings ( bytes ). However, Unicode strings and 8-bit strings cannot be mixed: that is, you cannot match a Unicode string with a byte pattern or vice-versa; similarly, when asking for a substitution, the replacement string must be of the same type as both the pattern and the search string.

Regular expressions use the backslash character ( ‘\’ ) to indicate special forms or to allow special characters to be used without invoking their special meaning. This collides with Python’s usage of the same character for the same purpose in string literals; for example, to match a literal backslash, one might have to write ‘\\\\’ as the pattern string, because the regular expression must be \\ , and each backslash must be expressed as \\ inside a regular Python string literal. Also, please note that any invalid escape sequences in Python’s usage of the backslash in string literals now generate a DeprecationWarning and in the future this will become a SyntaxError . This behaviour will happen even if it is a valid escape sequence for a regular expression.

The solution is to use Python’s raw string notation for regular expression patterns; backslashes are not handled in any special way in a string literal prefixed with ‘r’ . So r"\n" is a two-character string containing ‘\’ and ‘n’ , while "\n" is a one-character string containing a newline. Usually patterns will be expressed in Python code using this raw string notation.

It is important to note that most regular expression operations are available as module-level functions and methods on compiled regular expressions . The functions are shortcuts that don’t require you to compile a regex object first, but miss some fine-tuning parameters.

The third-party regex module, which has an API compatible with the standard library re module, but offers additional functionality and a more thorough Unicode support.

Regular Expression Syntax¶

A regular expression (or RE) specifies a set of strings that matches it; the functions in this module let you check if a particular string matches a given regular expression (or if a given regular expression matches a particular string, which comes down to the same thing).

Regular expressions can be concatenated to form new regular expressions; if A and B are both regular expressions, then AB is also a regular expression. In general, if a string p matches A and another string q matches B, the string pq will match AB. This holds unless A or B contain low precedence operations; boundary conditions between A and B; or have numbered group references. Thus, complex expressions can easily be constructed from simpler primitive expressions like the ones described here. For details of the theory and implementation of regular expressions, consult the Friedl book [Frie09] , or almost any textbook about compiler construction.

A brief explanation of the format of regular expressions follows. For further information and a gentler presentation, consult the Regular Expression HOWTO .

Regular expressions can contain both special and ordinary characters. Most ordinary characters, like ‘A’ , ‘a’ , or ‘0’ , are the simplest regular expressions; they simply match themselves. You can concatenate ordinary characters, so last matches the string ‘last’ . (In the rest of this section, we’ll write RE’s in this special style , usually without quotes, and strings to be matched ‘in single quotes’ .)

Some characters, like ‘|’ or ‘(‘ , are special. Special characters either stand for classes of ordinary characters, or affect how the regular expressions around them are interpreted.

Repetition operators or quantifiers ( * , + , ? , , etc) cannot be directly nested. This avoids ambiguity with the non-greedy modifier suffix ? , and with other modifiers in other implementations. To apply a second repetition to an inner repetition, parentheses may be used. For example, the expression (?:a<6>)* matches any multiple of six ‘a’ characters.

The special characters are:

(Dot.) In the default mode, this matches any character except a newline. If the DOTALL flag has been specified, this matches any character including a newline.

(Caret.) Matches the start of the string, and in MULTILINE mode also matches immediately after each newline.

Matches the end of the string or just before the newline at the end of the string, and in MULTILINE mode also matches before a newline. foo matches both ‘foo’ and ‘foobar’, while the regular expression foo$ matches only ‘foo’. More interestingly, searching for foo.$ in ‘foo1\nfoo2\n’ matches ‘foo2’ normally, but ‘foo1’ in MULTILINE mode; searching for a single $ in ‘foo\n’ will find two (empty) matches: one just before the newline, and one at the end of the string.

Causes the resulting RE to match 0 or more repetitions of the preceding RE, as many repetitions as are possible. ab* will match ‘a’, ‘ab’, or ‘a’ followed by any number of ‘b’s.

Causes the resulting RE to match 1 or more repetitions of the preceding RE. ab+ will match ‘a’ followed by any non-zero number of ‘b’s; it will not match just ‘a’.

Causes the resulting RE to match 0 or 1 repetitions of the preceding RE. ab? will match either ‘a’ or ‘ab’.

The ‘*’ , ‘+’ , and ‘?’ quantifiers are all greedy; they match as much text as possible. Sometimes this behaviour isn’t desired; if the RE <.*> is matched against ‘<a> b <c>’ , it will match the entire string, and not just ‘<a>’ . Adding ? after the quantifier makes it perform the match in non-greedy or minimal fashion; as few characters as possible will be matched. Using the RE <.*?> will match only ‘<a>’ .

Like the ‘*’ , ‘+’ , and ‘?’ quantifiers, those where ‘+’ is appended also match as many times as possible. However, unlike the true greedy quantifiers, these do not allow back-tracking when the expression following it fails to match. These are known as possessive quantifiers. For example, a*a will match ‘aaaa’ because the a* will match all 4 ‘a’ s, but, when the final ‘a’ is encountered, the expression is backtracked so that in the end the a* ends up matching 3 ‘a’ s total, and the fourth ‘a’ is matched by the final ‘a’ . However, when a*+a is used to match ‘aaaa’ , the a*+ will match all 4 ‘a’ , but when the final ‘a’ fails to find any more characters to match, the expression cannot be backtracked and will thus fail to match. x*+ , x++ and x?+ are equivalent to (?>x*) , (?>x+) and (?>x?) correspondingly.

New in version 3.11.

Specifies that exactly m copies of the previous RE should be matched; fewer matches cause the entire RE not to match. For example, a <6>will match exactly six ‘a’ characters, but not five.

Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as many repetitions as possible. For example, a <3,5>will match from 3 to 5 ‘a’ characters. Omitting m specifies a lower bound of zero, and omitting n specifies an infinite upper bound. As an example, a<4,>b will match ‘aaaab’ or a thousand ‘a’ characters followed by a ‘b’ , but not ‘aaab’ . The comma may not be omitted or the modifier would be confused with the previously described form.

Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as few repetitions as possible. This is the non-greedy version of the previous quantifier. For example, on the 6-character string ‘aaaaaa’ , a <3,5>will match 5 ‘a’ characters, while a<3,5>? will only match 3 characters.

Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as many repetitions as possible without establishing any backtracking points. This is the possessive version of the quantifier above. For example, on the 6-character string ‘aaaaaa’ , a<3,5>+aa attempt to match 5 ‘a’ characters, then, requiring 2 more ‘a’ s, will need more characters than available and thus fail, while a<3,5>aa will match with a <3,5>capturing 5, then 4 ‘a’ s by backtracking and then the final 2 ‘a’ s are matched by the final aa in the pattern. x+ is equivalent to (?>x) .

New in version 3.11.

Either escapes special characters (permitting you to match characters like ‘*’ , ‘?’ , and so forth), or signals a special sequence; special sequences are discussed below.

If you’re not using a raw string to express the pattern, remember that Python also uses the backslash as an escape sequence in string literals; if the escape sequence isn’t recognized by Python’s parser, the backslash and subsequent character are included in the resulting string. However, if Python would recognize the resulting sequence, the backslash should be repeated twice. This is complicated and hard to understand, so it’s highly recommended that you use raw strings for all but the simplest expressions.

Used to indicate a set of characters. In a set:

Characters can be listed individually, e.g. [amk] will match ‘a’ , ‘m’ , or ‘k’ .

Ranges of characters can be indicated by giving two characters and separating them by a ‘-‘ , for example [a-z] will match any lowercase ASCII letter, [0-5][0-9] will match all the two-digits numbers from 00 to 59 , and [0-9A-Fa-f] will match any hexadecimal digit. If — is escaped (e.g. [a\-z] ) or if it’s placed as the first or last character (e.g. [-a] or [a-] ), it will match a literal ‘-‘ .

Special characters lose their special meaning inside sets. For example, [(+*)] will match any of the literal characters ‘(‘ , ‘+’ , ‘*’ , or ‘)’ .

Character classes such as \w or \S (defined below) are also accepted inside a set, although the characters they match depends on whether ASCII or LOCALE mode is in force.

Characters that are not within a range can be matched by complementing the set. If the first character of the set is ‘^’ , all the characters that are not in the set will be matched. For example, [^5] will match any character except ‘5’ , and [^^] will match any character except ‘^’ . ^ has no special meaning if it’s not the first character in the set.

To match a literal ‘]’ inside a set, precede it with a backslash, or place it at the beginning of the set. For example, both [()[\]<>] and []()[<>] will match a right bracket, as well as left bracket, braces, and parentheses.

Support of nested sets and set operations as in Unicode Technical Standard #18 might be added in the future. This would change the syntax, so to facilitate this change a FutureWarning will be raised in ambiguous cases for the time being. That includes sets starting with a literal ‘[‘ or containing literal character sequences ‘—‘ , ‘&&’ , ‘

‘ , and ‘||’ . To avoid a warning escape them with a backslash.

Changed in version 3.7: FutureWarning is raised if a character set contains constructs that will change semantically in the future.

A|B , where A and B can be arbitrary REs, creates a regular expression that will match either A or B. An arbitrary number of REs can be separated by the ‘|’ in this way. This can be used inside groups (see below) as well. As the target string is scanned, REs separated by ‘|’ are tried from left to right. When one pattern completely matches, that branch is accepted. This means that once A matches, B will not be tested further, even if it would produce a longer overall match. In other words, the ‘|’ operator is never greedy. To match a literal ‘|’ , use \| , or enclose it inside a character class, as in [|] .

Matches whatever regular expression is inside the parentheses, and indicates the start and end of a group; the contents of a group can be retrieved after a match has been performed, and can be matched later in the string with the \number special sequence, described below. To match the literals ‘(‘ or ‘)’ , use \( or \) , or enclose them inside a character class: [(] , [)] .

This is an extension notation (a ‘?’ following a ‘(‘ is not meaningful otherwise). The first character after the ‘?’ determines what the meaning and further syntax of the construct is. Extensions usually do not create a new group; (?P<name>. ) is the only exception to this rule. Following are the currently supported extensions.

(One or more letters from the set ‘a’ , ‘i’ , ‘L’ , ‘m’ , ‘s’ , ‘u’ , ‘x’ .) The group matches the empty string; the letters set the corresponding flags: re.A (ASCII-only matching), re.I (ignore case), re.L (locale dependent), re.M (multi-line), re.S (dot matches all), re.U (Unicode matching), and re.X (verbose), for the entire regular expression. (The flags are described in Module Contents .) This is useful if you wish to include the flags as part of the regular expression, instead of passing a flag argument to the re.compile() function. Flags should be used first in the expression string.

Changed in version 3.11: This construction can only be used at the start of the expression.

A non-capturing version of regular parentheses. Matches whatever regular expression is inside the parentheses, but the substring matched by the group cannot be retrieved after performing a match or referenced later in the pattern.

(Zero or more letters from the set ‘a’ , ‘i’ , ‘L’ , ‘m’ , ‘s’ , ‘u’ , ‘x’ , optionally followed by ‘-‘ followed by one or more letters from the ‘i’ , ‘m’ , ‘s’ , ‘x’ .) The letters set or remove the corresponding flags: re.A (ASCII-only matching), re.I (ignore case), re.L (locale dependent), re.M (multi-line), re.S (dot matches all), re.U (Unicode matching), and re.X (verbose), for the part of the expression. (The flags are described in Module Contents .)

The letters ‘a’ , ‘L’ and ‘u’ are mutually exclusive when used as inline flags, so they can’t be combined or follow ‘-‘ . Instead, when one of them appears in an inline group, it overrides the matching mode in the enclosing group. In Unicode patterns (?a. ) switches to ASCII-only matching, and (?u. ) switches to Unicode matching (default). In byte pattern (?L. ) switches to locale depending matching, and (?a. ) switches to ASCII-only matching (default). This override is only in effect for the narrow inline group, and the original matching mode is restored outside of the group.

New in version 3.6.

Changed in version 3.7: The letters ‘a’ , ‘L’ and ‘u’ also can be used in a group.

Attempts to match . as if it was a separate regular expression, and if successful, continues to match the rest of the pattern following it. If the subsequent pattern fails to match, the stack can only be unwound to a point before the (?>. ) because once exited, the expression, known as an atomic group, has thrown away all stack points within itself. Thus, (?>.*). would never match anything because first the .* would match all characters possible, then, having nothing left to match, the final . would fail to match. Since there are no stack points saved in the Atomic Group, and there is no stack point before it, the entire expression would thus fail to match.

New in version 3.11.

Similar to regular parentheses, but the substring matched by the group is accessible via the symbolic group name name. Group names must be valid Python identifiers, and each group name must be defined only once within a regular expression. A symbolic group is also a numbered group, just as if the group were not named.

Regular Expressions: Regexes in Python (Part 2)

In the previous tutorial in this series, you covered a lot of ground. You saw how to use re.search() to perform pattern matching with regexes in Python and learned about the many regex metacharacters and parsing flags that you can use to fine-tune your pattern-matching capabilities.

But as great as all that is, the re module has much more to offer.

In this tutorial, you’ll:

  • Explore more functions, beyond re.search() , that the re module provides
  • Learn when and how to precompile a regex in Python into a regular expression object
  • Discover useful things that you can do with the match object returned by the functions in the re module

Ready? Let’s dig in!

Free Bonus: Get a sample chapter from Python Basics: A Practical Introduction to Python 3 to see how you can go from beginner to intermediate in Python with a complete curriculum, up to date for Python 3.9.

re Module Functions

In addition to re.search() , the re module contains several other functions to help you perform regex-related tasks.

Note: You saw in the previous tutorial that re.search() can take an optional <flags> argument, which specifies flags that modify parsing behavior. All the functions shown below, with the exception of re.escape() , support the <flags> argument in the same way.

You can specify <flags> as either a positional argument or a keyword argument:

The default for <flags> is always 0 , which indicates no special modification of matching behavior. Remember from the discussion of flags in the previous tutorial that the re.UNICODE flag is always set by default.

The available regex functions in the Python re module fall into the following three categories:

  1. Searching functions
  2. Substitution functions
  3. Utility functions

The following sections explain these functions in more detail.

Searching Functions

Searching functions scan a search string for one or more matches of the specified regex:

Function Description
re.search() Scans a string for a regex match
re.match() Looks for a regex match at the beginning of a string
re.fullmatch() Looks for a regex match on an entire string
re.findall() Returns a list of all regex matches in a string
re.finditer() Returns an iterator that yields regex matches from a string

As you can see from the table, these functions are similar to one another. But each one tweaks the searching functionality in its own way.

re.search(<regex>, <string>, flags=0)

Scans a string for a regex match.

If you worked through the previous tutorial in this series, then you should be well familiar with this function by now. re.search(<regex>, <string>) looks for any location in <string> where <regex> matches:

The function returns a match object if it finds a match and None otherwise.

Looks for a regex match at the beginning of a string.

This is identical to re.search() , except that re.search() returns a match if <regex> matches anywhere in <string> , whereas re.match() returns a match only if <regex> matches at the beginning of <string> :

In the above example, re.search() matches when the digits are both at the beginning of the string and in the middle, but re.match() matches only when the digits are at the beginning.

Remember from the previous tutorial in this series that if <string> contains embedded newlines, then the MULTILINE flag causes re.search() to match the caret ( ^ ) anchor metacharacter either at the beginning of <string> or at the beginning of any line contained within <string> :

The MULTILINE flag does not affect re.match() in this way:

Even with the MULTILINE flag set, re.match() will match the caret ( ^ ) anchor only at the beginning of <string> , not at the beginning of lines contained within <string> .

Note that, although it illustrates the point, the caret ( ^ ) anchor on line 3 in the above example is redundant. With re.match() , matches are essentially always anchored at the beginning of the string.

re.fullmatch(<regex>, <string>, flags=0)

Looks for a regex match on an entire string.

This is similar to re.search() and re.match() , but re.fullmatch() returns a match only if <regex> matches <string> in its entirety:

In the call on line 7, the search string ‘123’ consists entirely of digits from beginning to end. So that is the only case in which re.fullmatch() returns a match.

The re.search() call on line 10, in which the \d+ regex is explicitly anchored at the start and end of the search string, is functionally equivalent.

re.findall(<regex>, <string>, flags=0)

Returns a list of all matches of a regex in a string.

re.findall(<regex>, <string>) returns a list of all non-overlapping matches of <regex> in <string> . It scans the search string from left to right and returns all matches in the order found:

If <regex> contains a capturing group, then the return list contains only contents of the group, not the entire match:

In this case, the specified regex is #(\w+)# . The matching strings are ‘#foo#’ , ‘#bar#’ , and ‘#baz#’ . But the hash ( # ) characters don’t appear in the return list because they’re outside the grouping parentheses.

If <regex> contains more than one capturing group, then re.findall() returns a list of tuples containing the captured groups. The length of each tuple is equal to the number of groups specified:

In the above example, the regex on line 1 contains two capturing groups, so re.findall() returns a list of three two-tuples, each containing two captured matches. Line 4 contains three groups, so the return value is a list of two three-tuples.

re.finditer(<regex>, <string>, flags=0)

Returns an iterator that yields regex matches.

re.finditer(<regex>, <string>) scans <string> for non-overlapping matches of <regex> and returns an iterator that yields the match objects from any it finds. It scans the search string from left to right and returns matches in the order it finds them:

re.findall() and re.finditer() are very similar, but they differ in two respects:

re.findall() returns a list, whereas re.finditer() returns an iterator.

The items in the list that re.findall() returns are the actual matching strings, whereas the items yielded by the iterator that re.finditer() returns are match objects.

Any task that you could accomplish with one, you could probably also manage with the other. Which one you choose will depend on the circumstances. As you’ll see later in this tutorial, a lot of useful information can be obtained from a match object. If you need that information, then re.finditer() will probably be the better choice.

Substitution Functions

Substitution functions replace portions of a search string that match a specified regex:

Function Description
re.sub() Scans a string for regex matches, replaces the matching portions of the string with the specified replacement string, and returns the result
re.subn() Behaves just like re.sub() but also returns information regarding the number of substitutions made

Both re.sub() and re.subn() create a new string with the specified substitutions and return it. The original string remains unchanged. (Remember that strings are immutable in Python, so it wouldn’t be possible for these functions to modify the original string.)

re.sub(<regex>, <repl>, <string>, count=0, flags=0)

Returns a new string that results from performing replacements on a search string.

re.sub(<regex>, <repl>, <string>) finds the leftmost non-overlapping occurrences of <regex> in <string> , replaces each match as indicated by <repl> , and returns the result. <string> remains unchanged.

<repl> can be either a string or a function, as explained below.

Substitution by String

If <repl> is a string, then re.sub() inserts it into <string> in place of any sequences that match <regex> :

On line 3, the string ‘#’ replaces sequences of digits in s . On line 5, the string ‘(*)’ replaces sequences of lowercase letters. In both cases, re.sub() returns the modified string as it always does.

re.sub() replaces numbered backreferences ( \<n> ) in <repl> with the text of the corresponding captured group:

Here, captured groups 1 and 2 contain ‘foo’ and ‘qux’ . In the replacement string ‘\2,bar,baz,\1’ , ‘foo’ replaces \1 and ‘qux’ replaces \2 .

You can also refer to named backreferences created with (?P<name><regex>) in the replacement string using the metacharacter sequence \g<name> :

In fact, you can also refer to numbered backreferences this way by specifying the group number inside the angled brackets:

You may need to use this technique to avoid ambiguity in cases where a numbered backreference is immediately followed by a literal digit character. For example, suppose you have a string like ‘foo 123 bar’ and want to add a ‘0’ at the end of the digit sequence. You might try this:

Alas, the regex parser in Python interprets \10 as a backreference to the tenth captured group, which doesn’t exist in this case. Instead, you can use \g<1> to refer to the group:

The backreference \g<0> refers to the text of the entire match. This is valid even when there are no grouping parentheses in <regex> :

If <regex> specifies a zero-length match, then re.sub() will substitute <repl> into every character position in the string:

In the example above, the regex x* matches any zero-length sequence, so re.sub() inserts the replacement string at every character position in the string—before the first character, between each pair of characters, and after the last character.

If re.sub() doesn’t find any matches, then it always returns <string> unchanged.

Substitution by Function

If you specify <repl> as a function, then re.sub() calls that function for each match found. It passes each corresponding match object as an argument to the function to provide information about the match. The function return value then becomes the replacement string:

In this example, f() gets called for each match. As a result, re.sub() converts each alphanumeric portion of <string> to all uppercase and multiplies each numeric portion by 10 .

Limiting the Number of Replacements

If you specify a positive integer for the optional count parameter, then re.sub() performs at most that many replacements:

As with most re module functions, re.sub() accepts an optional <flags> argument as well.

re.subn(<regex>, <repl>, <string>, count=0, flags=0)

Returns a new string that results from performing replacements on a search string and also returns the number of substitutions made.

re.subn() is identical to re.sub() , except that re.subn() returns a two-tuple consisting of the modified string and the number of substitutions made:

In all other respects, re.subn() behaves just like re.sub() .

Utility Functions

There are two remaining regex functions in the Python re module that you’ve yet to cover:

Function Description
re.split() Splits a string into substrings using a regex as a delimiter
re.escape() Escapes characters in a regex

These are functions that involve regex matching but don’t clearly fall into either of the categories described above.

re.split(<regex>, <string>, maxsplit=0, flags=0)

Splits a string into substrings.

re.split(<regex>, <string>) splits <string> into substrings using <regex> as the delimiter and returns the substrings as a list.

The following example splits the specified string into substrings delimited by a comma ( , ), semicolon ( ; ), or slash ( / ) character, surrounded by any amount of whitespace:

If <regex> contains capturing groups, then the return list includes the matching delimiter strings as well:

This time, the return list contains not only the substrings ‘foo’ , ‘bar’ , ‘baz’ , and ‘qux’ but also several delimiter strings:

  • ‘,’
  • ‘ ; ‘
  • ‘ / ‘

This can be useful if you want to split <string> apart into delimited tokens, process the tokens in some way, then piece the string back together using the same delimiters that originally separated them:

If you need to use groups but don’t want the delimiters included in the return list, then you can use noncapturing groups:

If the optional maxsplit argument is present and greater than zero, then re.split() performs at most that many splits. The final element in the return list is the remainder of <string> after all the splits have occurred:

Explicitly specifying maxsplit=0 is equivalent to omitting it entirely. If maxsplit is negative, then re.split() returns <string> unchanged (in case you were looking for a rather elaborate way of doing nothing at all).

If <regex> contains capturing groups so that the return list includes delimiters, and <regex> matches the start of <string> , then re.split() places an empty string as the first element in the return list. Similarly, the last item in the return list is an empty string if <regex> matches the end of <string> :

In this case, the <regex> delimiter is a single slash ( / ) character. In a sense, then, there’s an empty string to the left of the first delimiter and to the right of the last one. So it makes sense that re.split() places empty strings as the first and last elements of the return list.

Escapes characters in a regex.

re.escape(<regex>) returns a copy of <regex> with each nonword character (anything other than a letter, digit, or underscore) preceded by a backslash.

This is useful if you’re calling one of the re module functions, and the <regex> you’re passing in has a lot of special characters that you want the parser to take literally instead of as metacharacters. It saves you the trouble of putting in all the backslash characters manually:

In this example, there isn’t a match on line 1 because the regex ‘foo^bar(baz)|qux’ contains special characters that behave as metacharacters. On line 3, they’re explicitly escaped with backslashes, so a match occurs. Lines 6 and 8 demonstrate that you can achieve the same effect using re.escape() .

Compiled Regex Objects in Python

The re module supports the capability to precompile a regex in Python into a regular expression object that can be repeatedly used later.

Compiles a regex into a regular expression object.

re.compile(<regex>) compiles <regex> and returns the corresponding regular expression object. If you include a <flags> value, then the corresponding flags apply to any searches performed with the object.

There are two ways to use a compiled regular expression object. You can specify it as the first argument to the re module functions in place of <regex> :

You can also invoke a method directly from a regular expression object:

Both of the examples above are equivalent to this:

Here’s one of the examples you saw previously, recast using a compiled regular expression object:

Here’s another, which also uses the IGNORECASE flag:

In this example, the statement on line 1 specifies regex ba[rz] directly to re.search() as the first argument. On line 4, the first argument to re.search() is the compiled regular expression object re_obj . On line 5, search() is invoked directly on re_obj . All three cases produce the same match.

Why Bother Compiling a Regex?

What good is precompiling? There are a couple of possible advantages.

If you use a particular regex in your Python code frequently, then precompiling allows you to separate out the regex definition from its uses. This enhances modularity. Consider this example:

Here, the regex \d+ appears several times. If, in the course of maintaining this code, you decide you need a different regex, then you’ll need to change it in each location. That’s not so bad in this small example because the uses are close to one another. But in a larger application, they might be widely scattered and difficult to track down.

The following is more modular and more maintainable:

Then again, you can achieve similar modularity without precompiling by using variable assignment:

In theory, you might expect precompilation to result in faster execution time as well. Suppose you call re.search() many thousands of times on the same regex. It might seem like compiling the regex once ahead of time would be more efficient than recompiling it each of the thousands of times it’s used.

In practice, though, that isn’t the case. The truth is that the re module compiles and caches a regex when it’s used in a function call. If the same regex is used subsequently in the same Python code, then it isn’t recompiled. The compiled value is fetched from cache instead. So the performance advantage is minimal.

All in all, there isn’t any immensely compelling reason to compile a regex in Python. Like much of Python, it’s just one more tool in your toolkit that you can use if you feel it will improve the readability or structure of your code.

Regular Expression Object Methods

A compiled regular expression object re_obj supports the following methods:

  • re_obj.search(<string>[, <pos>[, <endpos>]])
  • re_obj.match(<string>[, <pos>[, <endpos>]])
  • re_obj.fullmatch(<string>[, <pos>[, <endpos>]])
  • re_obj.findall(<string>[, <pos>[, <endpos>]])
  • re_obj.finditer(<string>[, <pos>[, <endpos>]])

These all behave the same way as the corresponding re functions that you’ve already encountered, with the exception that they also support the optional <pos> and <endpos> parameters. If these are present, then the search only applies to the portion of <string> indicated by <pos> and <endpos> , which act the same way as indices in slice notation:

In the above example, the regex is \d+ , a sequence of digit characters. The .search() call on line 4 searches all of s , so there’s a match. On line 9, the <pos> and <endpos> parameters effectively restrict the search to the substring starting with character 6 and going up to but not including character 9 (the substring ‘bar’ ), which doesn’t contain any digits.

If you specify <pos> but omit <endpos> , then the search applies to the substring from <pos> to the end of the string.

Note that anchors such as caret ( ^ ) and dollar sign ( $ ) still refer to the start and end of the entire string, not the substring determined by <pos> and <endpos> :

Here, even though ‘bar’ does occur at the start of the substring beginning at character 3, it isn’t at the start of the entire string, so the caret ( ^ ) anchor fails to match.

The following methods are available for a compiled regular expression object re_obj as well:

  • re_obj.split(<string>, maxsplit=0)
  • re_obj.sub(<repl>, <string>, count=0)
  • re_obj.subn(<repl>, <string>, count=0)

These also behave analogously to the corresponding re functions, but they don’t support the <pos> and <endpos> parameters.

Regular Expression Object Attributes

The re module defines several useful attributes for a compiled regular expression object:

Attribute Meaning
re_obj.flags Any <flags> that are in effect for the regex
re_obj.groups The number of capturing groups in the regex
re_obj.groupindex A dictionary mapping each symbolic group name defined by the (?P<name>) construct (if any) to the corresponding group number
re_obj.pattern The <regex> pattern that produced this object

The code below demonstrates some uses of these attributes:

Note that .flags includes any flags specified as arguments to re.compile() , any specified within the regex with the (?flags) metacharacter sequence, and any that are in effect by default. In the regular expression object defined on line 1, there are three flags defined:

  1. re.I : Specified as a <flags> value in the re.compile() call
  2. re.M : Specified as (?m) within the regex
  3. re.UNICODE : Enabled by default

You can see on line 4 that the value of re_obj.flags is the logical OR of these three values, which equals 42 .

The value of the .groupindex attribute for the regular expression object defined on line 11 is technically an object of type mappingproxy . For practical purposes, it functions like a dictionary.

Match Object Methods and Attributes

As you’ve seen, most functions and methods in the re module return a match object when there’s a successful match. Because a match object is truthy, you can use it in a conditional:

But match objects also contain quite a bit of handy information about the match. You’ve already seen some of it—the span= and match= data that the interpreter shows when it displays a match object. You can obtain much more from a match object using its methods and attributes.

Match Object Methods

The table below summarizes the methods that are available for a match object match :

Method Returns
match.group() The specified captured group or groups from match
match.__getitem__() A captured group from match
match.groups() All the captured groups from match
match.groupdict() A dictionary of named captured groups from match
match.expand() The result of performing backreference substitutions from match
match.start() The starting index of match
match.end() The ending index of match
match.span() Both the starting and ending indices of match as a tuple

The following sections describe these methods in more detail.

Returns the specified captured group(s) from a match.

For numbered groups, match.group(n) returns the n th group:

Remember: Numbered captured groups are one-based, not zero-based.

If you capture groups using (?P<name><regex>) , then match.group(<name>) returns the corresponding named group:

With more than one argument, .group() returns a tuple of all the groups specified. A given group can appear multiple times, and you can specify any captured groups in any order:

If you specify a group that’s out of range or nonexistent, then .group() raises an IndexError exception:

It’s possible for a regex in Python to match as a whole but to contain a group that doesn’t participate in the match. In that case, .group() returns None for the nonparticipating group. Consider this example:

This regex matches, as you can see from the match object. The first two captured groups contain ‘foo’ and ‘bar’ , respectively.

A question mark ( ? ) quantifier metacharacter follows the third group, though, so that group is optional. A match will occur if there’s a third sequence of word characters following the second comma ( , ) but also if there isn’t.

In this case, there isn’t. So there is match overall, but the third group doesn’t participate in it. As a result, m.group(3) is still defined and is a valid reference, but it returns None :

It can also happen that a group participates in the overall match multiple times. If you call .group() for that group number, then it returns only the part of the search string that matched the last time. The earlier matches aren’t accessible:

In this example, the full match is ‘foo,bar,baz,’ , as shown by the displayed match object. Each of ‘foo,’ , ‘bar,’ , and ‘baz,’ matches what’s inside the group, but m.group(1) returns only the last match, ‘baz,’ .

If you call .group() with an argument of 0 or no argument at all, then it returns the entire match:

This is the same data the interpreter shows following match= when it displays the match object, as you can see on line 3 above.

Returns a captured group from a match.

match.__getitem__(<grp>) is identical to match.group(<grp>) and returns the single group specified by <grp> :

If .__getitem__() simply replicates the functionality of .group() , then why would you use it? You probably wouldn’t directly, but you might indirectly. Read on to see why.

A Brief Introduction to Magic Methods

.__getitem__() is one of a collection of methods in Python called magic methods. These are special methods that the interpreter calls when a Python statement contains specific corresponding syntactical elements.

Note: Magic methods are also referred to as dunder methods because of the double underscore at the beginning and end of the method name.

Later in this series, there are several tutorials on object-oriented programming. You’ll learn much more about magic methods there.

The particular syntax that .__getitem__() corresponds to is indexing with square brackets. For any object obj , whenever you use the expression obj[n] , behind the scenes Python quietly translates it to a call to .__getitem__() . The following expressions are effectively equivalent:

The syntax obj[n] is only meaningful if a .__getitem()__ method exists for the class or type to which obj belongs. Exactly how Python interprets obj[n] will then depend on the implementation of .__getitem__() for that class.

Back to Match Objects

As of Python version 3.6, the re module does implement .__getitem__() for match objects. The implementation is such that match.__getitem__(n) is the same as match.group(n) .

The result of all this is that, instead of calling .group() directly, you can access captured groups from a match object using square-bracket indexing syntax instead:

This works with named captured groups as well:

This is something you could achieve by just calling .group() explicitly, but it’s a pretty shortcut notation nonetheless.

When a programming language provides alternate syntax that isn’t strictly necessary but allows for the expression of something in a cleaner, easier-to-read way, it’s called syntactic sugar. For a match object, match[n] is syntactic sugar for match.group(n) .

Note: Many objects in Python have a .__getitem()__ method defined, allowing the use of square-bracket indexing syntax. However, this feature is only available for regex match objects in Python version 3.6 or later.

Returns all captured groups from a match.

match.groups() returns a tuple of all captured groups:

As you saw previously, when a group in a regex in Python doesn’t participate in the overall match, .group() returns None for that group. By default, .groups() does likewise.

If you want .groups() to return something else in this situation, then you can use the default keyword argument:

Here, the third (\w+) group doesn’t participate in the match because the question mark ( ? ) metacharacter makes it optional, and the string ‘foo,bar,’ doesn’t contain a third sequence of word characters. By default, m.groups() returns None for the third group, as shown on line 8. On line 10, you can see that specifying default=’—‘ causes it to return the string ‘—‘ instead.

There isn’t any corresponding default keyword for .group() . It always returns None for nonparticipating groups.

Returns a dictionary of named captured groups.

match.groupdict() returns a dictionary of all named groups captured with the (?P<name><regex>) metacharacter sequence. The dictionary keys are the group names and the dictionary values are the corresponding group values:

As with .groups() , for .groupdict() the default argument determines the return value for nonparticipating groups:

Again, the final group (?P<w2>\w+) doesn’t participate in the overall match because of the question mark ( ? ) metacharacter. By default, m.groupdict() returns None for this group, but you can change it with the default argument.

Performs backreference substitutions from a match.

match.expand(<template>) returns the string that results from performing backreference substitution on <template> exactly as re.sub() would do:

This works for numeric backreferences, as on lines 7 and 9 above, and also for named backreferences, as on line 18.

Return the starting and ending indices of the match.

match.start() returns the index in the search string where the match begins, and match.end() returns the index immediately after where the match ends:

When Python displays a match object, these are the values listed with the span= keyword, as shown on line 4 above. They behave like string-slicing values, so if you use them to slice the original search string, then you should get the matching substring:

match.start(<grp>) and match.end(<grp>) return the starting and ending indices of the substring matched by <grp> , which may be a numbered or named group:

If the specified group matches a null string, then .start() and .end() are equal:

This makes sense if you remember that .start() and .end() act like slicing indices. Any string slice where the beginning and ending indices are equal will always be an empty string.

A special case occurs when the regex contains a group that doesn’t participate in the match:

As you’ve seen previously, in this case the third group doesn’t participate. m.start(3) and m.end(3) aren’t really meaningful here, so they return -1 .

Returns both the starting and ending indices of the match.

match.span() returns both the starting and ending indices of the match as a tuple. If you specified <grp> , then the return tuple applies to the given group:

The following are effectively equivalent:

  • match.span(<grp>)
  • (match.start(<grp>), match.end(<grp>))

match.span() just provides a convenient way to obtain both match.start() and match.end() in one method call.

Match Object Attributes

Like a compiled regular expression object, a match object also has several useful attributes available:

Attribute Meaning
match.pos
match.endpos
The effective values of the <pos> and <endpos> arguments for the match
match.lastindex The index of the last captured group
match.lastgroup The name of the last captured group
match.re The compiled regular expression object for the match
match.string The search string for the match

The following sections provide more detail on these match object attributes.

Contain the effective values of <pos> and <endpos> for the search.

Remember that some methods, when invoked on a compiled regex, accept optional <pos> and <endpos> arguments that limit the search to a portion of the specified search string. These values are accessible from the match object with the .pos and .endpos attributes:

If the <pos> and <endpos> arguments aren’t included in the call, either because they were omitted or because the function in question doesn’t accept them, then the .pos and .endpos attributes effectively indicate the start and end of the string:

The re_obj.search() call above on line 2 could take <pos> and <endpos> arguments, but they aren’t specified. The re.search() call on line 8 can’t take them at all. In either case, m.pos and m.endpos are 0 and 9 , the starting and ending indices of the search string ‘foo123bar’ .

Contains the index of the last captured group.

match.lastindex is equal to the integer index of the last captured group:

In cases where the regex contains potentially nonparticipating groups, this allows you to determine how many groups actually participated in the match:

In the first example, the third group, which is optional because of the question mark ( ? ) metacharacter, does participate in the match. But in the second example it doesn’t. You can tell because m.lastindex is 3 in the first case and 2 in the second.

There’s a subtle point to be aware of regarding .lastindex . It isn’t always the case that the last group to match is also the last group encountered syntactically. The Python documentation gives this example:

The outermost group is ((a)(b)) , which matches ‘ab’ . This is the first group the parser encounters, so it becomes group 1. But it’s also the last group to match, which is why m.lastindex is 1 .

The second and third groups the parser recognizes are (a) and (b) . These are groups 2 and 3 , but they match before group 1 does.

Contains the name of the last captured group.

If the last captured group originates from the (?P<name><regex>) metacharacter sequence, then match.lastgroup returns the name of that group:

match.lastgroup returns None if the last captured group isn’t a named group:

As shown above, this can be either because the last captured group isn’t a named group or because there were no captured groups at all.

Contains the regular expression object for the match.

match.re contains the regular expression object that produced the match. This is the same object you’d get if you passed the regex to re.compile() :

Remember from earlier that the re module caches regular expressions after it compiles them, so they don’t need to be recompiled if used again. For that reason, as the identity comparisons on lines 12 and 20 show, all the various regular expression objects in the above example are the exact same object.

Once you have access to the regular expression object for the match, all of that object’s attributes are available as well:

Here, .match() is invoked on m.re to perform another search using the same regex but on a different search string.

Contains the search string for a match.

match.string contains the search string that is the target of the match:

As you can see from the example, the .string attribute is available when the match object derives from a compiled regular expression object as well.

Conclusion

That concludes your tour of Python’s re module!

This introductory series contains two tutorials on regular expression processing in Python. If you’ve worked through both the previous tutorial and this one, then you should now know how to:

  • Make full use of all the functions that the re module provides
  • Precompile a regex in Python
  • Extract information from match objects

Regular expressions are extremely versatile and powerful—literally a language in their own right. You’ll find them invaluable in your Python coding.

Note: The re module is great, and it will likely serve you well in most circumstances. However, there’s an alternative third-party Python module called regex that provides even greater regular expression matching capability. You can learn more about it at the regex project page.

Next up in this series, you’ll explore how Python avoids conflict between identifiers in different areas of code. As you’ve already seen, each function in Python has its own namespace, distinct from those of other functions. In the next tutorial, you’ll learn how namespaces are implemented in Python and how they define variable scope.

re — Операции с регулярными выражениями¶

Модуль предоставляет операции сопоставления регулярных выражений, аналогичные тем, что есть в Perl.

Шаблоны и строки для поиска могут быть строками Юникода ( str ) а также 8-битными строками ( bytes ). Однако Юникод строки и 8-битные строки нельзя смешивать: то есть вы не можете сопоставить строку Юникод с байтовым шаблоном или наоборот; аналогично, при запросе замены, замена строка должна быть того же типа, что и шаблона, и строкой поиска.

Регулярные выражения используют символ обратной косой черты ( ‘\’ ), чтобы указать на специальные формы или позволить специальным знакам быть используемый, не призывая их специальное значение. Это сталкивается с использованием Python’s того же символ для той же цели в опечатках строка; например, чтобы соответствовать обратной косой черте литерал, возможно, придется написать ‘\\\\’ как шаблон строка, потому что регулярное выражение должно быть \\ , и каждая обратная косая черта должна быть выражена как \\ в регулярном Python строка литерал. Кроме того, обратите внимание на то, что любые недействительные последовательности экранирование в использовании Python’s обратной косой черты в опечатках строка теперь производят DeprecationWarning , и в будущем это станет SyntaxError . Это поведение произойдет даже в том случае, если это допустимая escape-последовательность для регулярного выражения.

Решение состоит в том, чтобы использовать сырое примечание строка Python’s для шаблонов регулярного выражения; обратная косая черта не обрабатывается каким-либо специальным образом в строковом литерале с префиксом ‘r’ . Таким образом, r"\n" — двухсимвольная строка, содержащая ‘\’ и ‘n’ , в то время как "\n" — одиносимвольная строка, определяющая новую строку. Обычно шаблоны будут выражаться в Python код с помощью этой необработанной нотации строки.

Важно отметить, что операции по наиболее регулярному выражению доступны, поскольку уровень модуля функционирует и методы на скомпилированные регулярные выражения . Функции — это ярлыки, которые не требуют сначала компиляции объекта regex, но пропускают некоторые параметры точной настройки.

Сторонний модуль regex, который имеет API, совместимый со стандартным модулем библиотеки re , но предлагает дополнительную функциональность и более полную поддержку Юникод.

Синтаксис регулярного выражения¶

Регулярное выражение (или RE) указывает множество строк, соответствующий ему; функции в этом модуле позволяют вам проверить, соответствует ли особый строка данному регулярному выражению (или если данное регулярное выражение соответствует особому строка, который сводится к тому же самому).

Регулярные выражения могут быть объединены для формирования новых регулярных выражений; если A и B оба являются регулярными выражениями, то AB также является регулярным выражением. В общем, если строка p совпадает A и ещё строка q совпадает B, строка pq будет соответствовать AB. Это сохраняется, если только A или B не содержат операций с низким приоритетом; граничные условия между A и B; или иметь нумерованные ссылки на группы. Таким образом, сложные выражения можно легко построить из более простых примитивных выражений, подобных описанным здесь. Подробности теории и реализации регулярных выражений см. в книге Friedl [Frie09], или почти любом учебнике о построении компилятора.

Ниже приводится краткое объяснение формата регулярных выражений. Для получения дополнительной информации и более мягкой презентации обратитесь к HOWTO по регулярным выражениям .

Регулярные выражения могут содержать как специальные, так и обычные символы. Большинство обычных символов, таких как ‘A’ , ‘a’ или ‘0’ , являются простейшими регулярными выражениями; они просто совпадают друг с другом. Можно объединять обычные символы, поэтому last соответствует строка ‘last’ . (В остальной части этого раздела мы запишем RE в особого стиля , обычно без кавычек, и строки для соответствия ‘в одинарных кавычках’ .

Некоторые символы, как ‘|’ или ‘(‘ , особенные. Специальные символы либо стоят для классов обычных символов, либо влияют на то, как интерпретируются регулярные выражения вокруг них.

Квалификаторы повторения ( * , + , ? , и т.д.) не могут быть вложены напрямую. Это позволяет избежать неоднозначности с нежадным суффиксом модификатора ? и с другими модификаторами в других реализациях. Чтобы применить второе повторение к внутреннему повторению, скобки могут быть используемый. Например, выражение (?:a<6>)* соответствует любому множеству из шести символов ‘a’ .

. (Точка.) в режиме по умолчанию это соответствует любому символу, кроме новой строки. Если флаг DOTALL был указан, он соответствует любому символу, включая новую строку. ^ (Каретка) Соответствует началу строки, а в режиме MULTILINE также совпадает сразу после каждой новой строки. $ Соответствует концу строки или непосредственно перед новой строкой в конце строка, а в режиме MULTILINE также совпадает перед новой строкой. foo соответствует и „foo“ и „foobar“, в то время как регулярное выражение foo$ соответствует только „foo“. Более интересно, что поиск foo.$ в ‘foo1\nfoo2\n’ соответствует „foo2“ обычно, но „foo1“ в режиме MULTILINE ; поиск одного $ в ‘foo\n’ найдёт два (пустых) совпадения: одно непосредственно перед новой строке, и одно в конце строка. * Приводит к тому, что результирующее RE соответствует 0 или более повторениям предыдущего RE, как можно больше повторений. ab* соответствует «a», «ab» или «a», за которым следует любое число «b». + Приводит к совпадению результирующего RE с 1 или более повторениями предыдущего RE. ab+ совпадет с «a», за которым следует любое ненулевое число «b»; он не будет соответствовать только „a“. ? Результирующее RE соответствует 0 или 1 повторениям предыдущего RE. ab? будет совпадать с „a“ или „ab“. *? , +? , ?? Квалификаторы ‘*’ , ‘+’ и ‘?’ — все жадные; они соответствуют как можно большему количеству текста. Иногда такое поведение не желательно; если RE, <.*> соответствует ‘<a> b <c>’ , он будет соответствовать всей строке и не просто ‘<a>’ . Добавление ? после квалификатора заставляет его выполнять соотвествие нежаднымy или минимальному способу; будут сопоставлены как можно более несколько символов. Используя RE <.*?> будет соответствовать только ‘<a>’ . Указывает, что должны быть сопоставлены точно m копии предыдущего RE; меньшее количество совпадений приводит к несоответствию всего RE. Например, a <6>будет соответствовать ровно шести ‘a’ символам, но не пяти. Приводит к совпадению результирующего RE от m к n повторениям предыдущего RE, пытаясь сопоставить как можно больше повторений. Например, a <3,5>будет совпадать от 3 до 5 символов ‘a’ . Опущение m указывает нижнюю границу нуля, а опущение n указывает бесконечную верхнюю границу. В качестве примера a<4,>b будет соответствовать ‘aaaab’ или тысяче ‘a’ символов, за которыми следует ‘b’ , но не ‘aaab’ . Запятая не может быть опущена или модификатор будет смешан с ранее описанной формой. ? Приводит к совпадению результирующего RE от m к n повторениям предыдущего RE, пытаясь сопоставить как можно более few повторений. Это нежадная версия предыдущего квалификатора. Например, на 6-символ строка ‘aaaaaa’ a <3,5>будет совпадать с 5 ‘a’ символами, в то время как a<3,5>? будет совпадать только с 3 символами. \

Либо экранирует от специальных символов (позволяя сопоставлять символы типа ‘*’ , ‘?’ и так далее), либо сигнализирует о специальной последовательности; специальные последовательности обсуждаются ниже.

Если вы не используете сырую строку, чтобы выразить шаблон, помните, что Python также использует обратную косую черту в качестве последовательности экранирование в строковых литералах; если последовательность экранирование не распознается Python парсером, обратная косая черта и последующие символ включаются в результирующий строка. Однако если Python распознает результирующую последовательность, обратная косая черта должна повторяться дважды. Это сложно и трудно понять, таким образом, настоятельно рекомендовано, что вы используете сырой строки для всех кроме самых простых выражений.

Используется для обозначения множества символов. В множестве:

  • Символы могут быть перечислены по отдельности, например, [amk] будет соответствовать ‘a’ , ‘m’ или ‘k’ .
  • Диапазоны символов можно указать, выдав два символа и разделив их по ‘-‘ , например [a-z] будет соответствовать любой строчной букве ASCII, [0-5][0-9] будет соответствовать всем двухзначным числам от 00 до 59 , а [0-9A-Fa-f] будет соответствовать любой шестнадцатеричной цифре. Если — не удалось выполнить (например, [a\-z] ) или если он помещен в качестве первого или последнего символ (например, [-a] или [a-] ), он будет соответствовать литералу ‘-‘ .
  • Специальные символы теряют свой особый смысл внутри множеств. Например, [(+*)] будет соответствовать любому из знаков литерала ‘(‘ , ‘+’ , ‘*’ или ‘)’ .
  • Классы символов, такие как \w или \S (определенные ниже), также принимаются внутри множества, хотя совпадающие символы зависят от того, действует ли режим ASCII или LOCALE .
  • Знаки, которые не являются в диапазоне, могут соответствовать дополняя множество. Если первый символ аппарата равен ‘^’ , то все символы, не в множестве, будут сопоставлены. Например, [^5] будет соответствовать любому символ, кроме ‘5’ , и [^^] будет соответствовать любому символу, кроме ‘^’ . ^ не имеет особого значения, если он не первый символ в множестве.
  • Чтобы сопоставить литерал ‘]’ внутри множества, перед ним стоит обратная косая черта или поместите его в начале множества. Например, и [()[\]<>] , и []()[<>] будут совпадать с круглыми скобками.

Изменено в версии 3.7: Поднимается FutureWarning , если множество символов содержит конструкции, которые семантически изменятся в будущем.

(Ноль или больше символов от множества ‘a’ , ‘i’ , ‘L’ , ‘m’ , ‘s’ , ‘u’ , ‘x’ , произвольно сопровождаемый ‘-‘ , сопровождаемым одним или несколькими символами от ‘i’ , ‘m’ , ‘s’ , ‘x’ .) буквы задают или удаляют соответствующие флаги: re.A (сопоставление только ASCII), re.I (игнорирование регистра), re.L (зависимое от локаль), re.M (многострочное), re.S (совпадение точек со всеми), re.U (совпадение Юникода) и re.X (детальное), для части выражения. (Флаги описаны в Содержимое модуля .

Буквы ‘a’ , ‘L’ и ‘u’ являются взаимоисключающими, когда используемый являются встроенными флагами, поэтому они не могут быть объединены или следовать за ‘-‘ . Вместо этого, когда один из них появляется во встроенной группе, он переопределяет режим сопоставления в заключительной группе. В шаблонах Юникода (?a. ) переключается на сопоставление только в ASCII, а (?u. ) переключается на соответствие в юникоде (по умолчанию). В шаблоне байтов (?L. ) переключается на локаль в зависимости от соответствия, а (?a. ) переключается на соответствие только ASCII (по умолчанию). Это переопределение действует только для узкой встроенной группы, и исходный режим сопоставления восстанавливается вне группы.

Добавлено в версии 3.6.

Изменено в версии 3.7: Буквы ‘a’ , ‘L’ и ‘u’ также могут быть используемый в группе.

аналогично обычным скобкам, но подстрока, совпадающая с группой, доступна через символическое имя группы name. Названия группы должны быть действительными идентификаторами Python, и каждое название группы должно быть определено только однажды в регулярном выражении. Символическая группа — это также нумерованная группа, как если бы она не называлась.

На именованные группы можно ссылаться в трех контекстах. Если шаблон имеет значение (?P<quote>[‘"]).*?(?P=quote) (т.е. соответствует строка, котированному с одинарными или двойными кавычками):

  • (?P=quote) (как показано)
  • \1
  • m.group(‘quote’)
  • m.end(‘quote’) (и т.д.)
  • \g<quote>
  • \g<1>
  • \1

Соответствует, если текущей позиции в строки предшествует совпадение для . , которое заканчивается на текущей позиции. Это называется позитивное опережающее утверждение. (?<=abc)def найдет совпадение в ‘abcdef’ , так как опережающее будет резервировать 3 символа и проверять, соответствует ли содержащийся шаблон. Содержавший шаблон должен только соответствовать строки некоторой фиксированной длины, означая, что abc или a|b позволены, но a* и a <3,4>не. Обратите внимание, что шаблоны, которые начинаются с положительных опережающее утверждений, не будут соответствовать в начале строки, являющегося найденный; скорее всего, потребуется использовать функцию search() , а не функцию match() :

В этом примере выполняется поиск слова после дефиса:

Изменено в версии 3.5: Добавлена поддержка групповых ссылок фиксированной длины.

Специальные последовательности состоят из ‘\’ и символ из списка ниже. Если обычный символ не является цифрой ASCII или буквой ASCII, то результирующий RE будет соответствовать второму символ. Например, \$ соответствует символу ‘$’ .

\number Соответствует содержимому группы с тем же номером. Группы нумеруются, начиная с 1. Например, (.+) \1 соответствует ‘the the’ или ’55 55′ , но не ‘thethe’ (обратите внимание на пробел после группы). Эта специальная последовательность может быть используемый только для соответствия одной из первых 99 групп. Если первая цифра number будет 0, или number — 3 октальных цифры долго, то это не будет интерпретироваться как матч группы, но как символ с октальным значение number. Внутри ‘[‘ и ‘]’ класса символ все числовые экранирования обрабатываются как символы. \A совпадает только в начале строка. \b

Соответствует пустой строке, но только в начале или конце слова. Слово определяется как последовательность символов слова. Заметим, что формально \b определяется как граница между \w и \W символ (или наоборот), или между \w и началом/концом строка. Это означает, что r’\bfoo\b’ соответствует ‘foo’ , ‘foo.’ , ‘(foo)’ , ‘bar foo baz’ , но не ‘foobar’ или ‘foo3’ .

По умолчанию буквенно-цифровой индикатор Юникод — те используемый в шаблонах Юникод, но это может быть изменено при помощи флага ASCII . Границы слов определяются текущим значением локали, если флаг LOCALE имеет значение используемый. Внутри диапазона символ \b представляет символ заднего пространства для совместимости с литералами Python’а строки.

\B соответствуют пустому строка, но только когда это — not вначале или конец слова. Это означает, что r’py\B’ соответствует ‘python’ , ‘py3’ , ‘py2’ , но не ‘py’ , ‘py.’ или ‘py!’ . \B является как раз противоположным \b , поэтому символы слов в шаблонах Юникода являются буквенно-цифровыми кодами Юникода или подчеркиванием, хотя это можно изменить с помощью флага ASCII . Границы слов определяются текущим значением локаль, если флаг LOCALE имеет значение используемый.

\d для Юникод (str) шаблоны: Соответствует любая десятичная цифра Юникод (то есть, любой символ в категории Юникод символ [Без обозначения даты]). Сюда входят [0-9] , а также множество других цифр. Если флаг ASCII — используемый, только [0-9] подобран. Для 8-разрядных (байт) шаблонов: соответствует любой десятичной цифре; это эквивалентно [0-9] . \D соответствует любой символ, которая не является десятичной цифрой. Это противоположность \d . Если флаг ASCII равен используемый, это становится эквивалентом [^0-9] . \s Для шаблонов Юникода (str): соответствует символам пробела Юникода (который включает в себя символы [ \t\n\r\f\v] , а также многие другие символы, например неразрывные пробелы, предписанные правилами типографики на многих языках). Если флаг ASCII равен используемый, сопоставляется только [ \t\n\r\f\v] . Для 8 битов (байт) шаблоны: Символы соответствия рассмотрели пробел в множестве ASCII символ; это эквивалентно [ \t\n\r\f\v] . \S Соответствует любому символ, который не является пробелом символ. Это противоположность \s . Если флаг ASCII равен используемый, это становится эквивалентом [^ \t\n\r\f\v] . \w Для шаблонов Юникода (str): Соответствует символам слова Юникода; это включает большинство символов, которые могут быть частью слова на любом языке, а также цифры и знак подчеркивания. Если флаг ASCII равен используемый, сопоставляется только [a-zA-Z0-9_] . Для 8 битных (байт) шаблонов: Соответствие символам считали алфавитно-цифровым в множестве ASCII символ; это эквивалентно [a-zA-Z0-9_] . Если флаг LOCALE равен используемый, совпадает с буквенно-цифровыми символами в текущей локали и подчеркиванием. \W Соответствует любому символу, который не является словом символ. Это противоположность \w . Если флаг ASCII равен используемый, это становится эквивалентом [^a-zA-Z0-9_] . Если флаг LOCALE равен используемый, соответствует символам, которые не являются ни буквенно- цифровыми в текущей локали, ни подчеркиванием. \Z Соответствует только концу строки.

Большинство стандартных экранирований, поддерживаемых Python строка литералами, также принимаются регулярным выражением парсер:

(Обратите внимание, что \b — используемый, чтобы представлять границы слова и означает «клавишу Backspace» только в классах символ.)

escape-последовательности ‘\u’ , ‘\U’ и ‘\N’ распознаются только в шаблонах Юникода. В шаблонах байтов это ошибки. Неизвестные экранирование символов ASCII зарезервированы для будущего использования и рассматриваются как ошибки.

Восьмеричные экранирование входят в ограниченном виде. Если первая цифра является 0, или если есть три восьмеричные цифры, она считается восьмеричным экранированием. В противном случае это ссылка на группу. Что касается строка литералов, восьмеричные экранирование всегда имеют длину максимум три цифры.

Изменено в версии 3.3: Последовательности экранирования ‘\u’ и ‘\U’ были добавлены.

Изменено в версии 3.6: Неизвестные экранирования, состоящие из ‘\’ и символа ASCII, теперь являются ошибками.

Изменено в версии 3.8: Последовательность экранирования ‘\N‘ была добавлена. Как в опечатках строка, это расширяется до названного Юникод символ (например, ‘\N‘ ).

Содержимое модуля¶

Модуль определяет несколько функций, констант и исключения. Некоторые из функций являются упрощенными версиями полнофункциональных методов для скомпилированных регулярных выражений. Большинство нетривиальных приложений всегда используют скомпилированную форму.

Изменено в версии 3.6: Константы флага теперь являются сущности RegexFlag , который является подкласс enum.IntFlag .

Скомпилировать шаблон регулярного выражения в объект регулярного выражения , который может быть используемый для сопоставления с помощью его match() , search() и других методов, описанных ниже.

Поведение выражения можно изменить, указав flags значение. Значения могут быть любыми из следующих переменных, объединяемых с помощью побитового оператора OR (оператор | ).

но использование re.compile() и сохранение результирующего объекта регулярного выражения для повторного использования более эффективно, когда выражение будет используемый несколько раз в одной программе.

Скомпилированные версии последних шаблонов, переданных re.compile() , и функции сопоставления на уровне модуля кэшируются, поэтому программам, которые используют только несколько регулярных выражений одновременно, не нужно беспокоиться о компиляции регулярных выражений.

Сделать \w , \W , \b , \B , \d , \D , \s и \S выполняют соответствие только для ASCII вместо полного соответствия Юникод. Это имеет значение только для шаблонов Юникода и игнорируется для шаблонов байтов. Соответствует встроенному флагу (?a) .

Обратите внимание, что для обратной совместимости флаг re.U все еще существует (а также его синоним re.Юникод и встроенный аналог (?u) ), но они являются избыточными в Python 3, так как совпадения по умолчанию являются юникодом для строки (и соответствие юникоду не допускается для байт).

Отображение отладочной информации о скомпилированном выражении. Соответствующий встроенный флаг отсутствует.

re. I ¶ re. IGNORECASE ¶

Выполнить согласование без учета регистра; выражения типа [A-Z] также будут соответствовать строчным буквам. Полное соответствие Юникод (такое как Ü , соответствующий ü ) также, работает, если флаг re.ASCII не используемый, чтобы отключить матчи неASCII. Текущий локали не изменяет эффект этого флага, если флаг re.LOCALE не также используемый. Соответствует встроенному флагу (?i) .

Обратите внимание, что когда шаблоны Юникода [a-z] или [A-Z] используемый в сочетании с флагом IGNORECASE , они будут соответствовать 52 буквам ASCII и 4 дополнительным буквам не ASCII: „İ“ (U+0130, латинская буква I с точкой выше), „ı“ (U+0131, латинская маленькая буква без точки i), „ſ“ (U+017F, латинская маленькая если флаг ASCII равен используемый, сопоставляются только буквы «a» — «z» и «A» — «Z».

re. L ¶ re. LOCALE ¶

Сделайте \w , \W , \b , \B и соответствие без учета регистра зависящему от текущей локали. Этот флаг может быть используемый только с шаблонами байтов. Использование этого флага не рекомендуется, так как механизм локали очень ненадежен, он обрабатывает только одну «культуру» за раз, и работает только с 8-битными локалями. Сопоставление Юникода уже включено по умолчанию в Python 3 для шаблонов Юникода (str), и он может обрабатывать различные локали/языки. Соответствует встроенному флагу (?L) .

Изменено в версии 3.6: re.LOCALE может быть используемый только с шаблонами байтов и несовместим с re.ASCII .

Изменено в версии 3.7: Скомпилированные объекты регулярного выражения с флагом re.LOCALE больше не зависят от локаль во время компиляции. Только локаль при соответствии времени затрагивает результат соответствия.

Когда он определен, символ ‘^’ образца соответствует в начале строка и в начале каждой линии (сразу после каждого newline); и символ ‘$’ образца соответствует в конце строка и в конце каждой линии (немедленно предшествующий каждому newline). По умолчанию ‘^’ совпадает только в начале строка, а ‘$’ только в конце строка и непосредственно перед новой линией (при ее наличии) в конце строка. Соответствует встроенному флагу (?m) .

re. S ¶ re. DOTALL ¶

Заставьте специальный символ ‘.’ соответствовать любому символ вообще, включая newline; без этого флага ‘.’ будет соответствовать чему- либо except новой строке. Соответствует встроенному флагу (?s) .

re. X ¶ re. VERBOSE ¶

Этот флаг позволяет писать регулярные выражения, которые выглядят красивее и более читаемы, позволяя визуально разделять логические разделы шаблона и добавлять комментарии. Пробельное пространство в шаблоне игнорируется, за исключением случаев, когда оно находится в классе символ или предваряется необъявленной обратной косой чертой, или в таких маркерах, как *? , (?: или (?P<. > . Когда линия содержит # , который не находится в классе символ и не предшествуется несбежавшей обратной косой чертой, всеми знаками от крайнего левого, такие # через конец линии проигнорированы.

Это означает, что два следующих объекта регулярного выражения, которые соответствуют десятичному числу, функционально равны:

Соответствует встроенному флагу (?x) .

Отсканировать string в поисках первого местоположения, в котором регулярное выражение pattern создает совпадение, и возвращает соответствующий объект соответствия . Возвращает None если ни одна позиция в строка не соответствует шаблону; следует отметить, что это отличается от поиска совпадения нулевой длины в какой-то точке строка.

Если ноль или больше знаков в начале string соответствуют регулярному выражению pattern, возвращает соответствующий объект соответствия . Возвращает None , если строка не соответствует шаблону; следует отметить, что это отличается от совпадения нулевой длины.

Отметим, что даже в режиме MULTILINE re.match() будет совпадать только в начале строка, а не в начале каждой строки.

Если требуется найти совпадение в любом месте string, используйте команду search() (см. также search() против match() ).

Если string соответствует целиком регулярному выражению pattern, возвращает соответствующий объект соответствия . Возвращает None , если строка не соответствует шаблону; следует отметить, что это отличается от совпадения нулевой длины.

Добавлено в версии 3.4.

Разделить string по вхождениям pattern. Если захват скобок используемый в pattern, то текст всех групп в шаблоне также возвращенный как часть результирующего списка. Если maxsplit отличный от нуля в большей части maxsplit, разделения происходят, и остаток от строка — возвращенный как заключительный элемент списка:

Если в разделителе есть группы захвата и он совпадает в начале строка, результат будет начинаться с пустого строка. То же самое относится и к концу строка:

Таким образом, компоненты разделителя всегда находятся в одном и том же относительном индексе в списке результатов.

Пустые совпадения для шаблона разделяют строка только в том случае, если они не примыкают к предыдущему пустому совпадению.

Изменено в версии 3.1: Добавлен аргумент дополнительных флагов.

Изменено в версии 3.7: Добавлена поддержка разделения на шаблоне, который может соответствовать пустой строка.

Возвращает все неперекрывающиеся совпадения pattern в string, как список строки. string сканируется слева направо, а совпадения возвращенный в найденном порядке. Если в шаблоне присутствует одна или несколько групп, возвращает список групп; это будет список кортежей, если шаблон имеет более одной группы. В результат включаются пустые совпадения.

Изменено в версии 3.7: Непустые совпадения теперь могут начинаться сразу после предыдущего пустого совпадения.

Возвращает итератор , приводящий к объектам соответствия по всему неперекрыванию, соответствует для RE pattern в string. string сканируется слева направо, а совпадения возвращенный в найденном порядке. В результат включаются пустые совпадения.

Изменено в версии 3.7: Непустые совпадения теперь могут начинаться сразу после предыдущего пустого совпадения.

Возвращает строку, полученные заменой крайних левых неперекрывающихся вхождений pattern в string заменяющим repl. Если шаблон не найден, string возвращенный неизменный. repl может быть строка или функцией; если это строка, то все экранирования обратной косой чертой в нем обрабатываются. То есть \n преобразуется в одну новую строку символ, \r преобразуется в каретку возвращает и так далее. Неизвестные экранирование символов ASCII зарезервированы для будущего использования и рассматриваются как ошибки. Другие неизвестные экранирования, такие как \& , остаются одни. Обратные ссылки, такие как \6 , заменяются подстрокой, соответствующей группе 6 в шаблоне. Например:

Если repl является функцией, она вызывается для каждого неперекрывающегося вхождения pattern. Функция принимает один аргумент объект соответствия и возвращает замену строка. Например:

Шаблон может быть строкой или объектом шаблона .

Необязательный аргумент count — максимальное число заменяемых экземпляров шаблона; count должно быть неотрицательным целым числом. Если значение опущено или равно нулю, все вхождения будут заменены. Пустые матчи для образца заменены только если не смежные с предыдущим пустым матчем, таким образом, sub(‘x*’, ‘-‘, ‘abxd’) возвращает ‘-a-b—d-‘ .

В аргументах строка-type repl, в дополнение к вышеописанным экранирование символ и обратным ссылкам, \g<name> будет использовать подстроку, соответствующую группе с именем name , как определено синтаксисом (?P<name>. ) . \g<number> использует соответствующий номер группы; поэтому \g<2> эквивалентен \2 , но не является неоднозначным в замене, такой как \g<2>0 . \20 будет интерпретироваться как ссылка на группу 20, а не как ссылка на группу 2, за которой следует группа литерал символ ‘0’ . Обратная ссылка \g<0> заменяет во всей подстроке, соответствующей RE.

Изменено в версии 3.1: Добавлен аргумент дополнительных флагов.

Изменено в версии 3.5: Несопоставленные группы заменяются пустым строка.

Изменено в версии 3.6: Неизвестные экранирование в pattern, состоящей из ‘\’ и символа ASCII, теперь являются ошибками.

Изменено в версии 3.7: Неизвестные экранирование в repl, состоящей из ‘\’ и символа ASCII, теперь являются ошибками.

Изменено в версии 3.7: Пустые совпадения для шаблона заменяются при соседстве с предыдущим непустым совпадением.

Выполнить ту же операцию, что и sub() , но возвращает кортеж (new_string, number_of_subs_made) .

Изменено в версии 3.1: Добавлен аргумент дополнительных флагов.

Изменено в версии 3.5: Несопоставленные группы заменяются пустым строка.

Экранирование специальных символов в pattern. Это полезно, если вы хотите соответствовать произвольному литерал строка, у которого могут быть метазнаки регулярного выражения в нем. Например:

Эта функция не должна быть используемый для замены строка в sub() и subn() , следует избегать только обратной косой черты. Например:

Изменено в версии 3.3: От ‘_’ символ больше не экранируется.

Изменено в версии 3.7: Ускользают только символы, которые могут иметь особое значение в регулярном выражении. В результате ‘!’ , ‘"’ , ‘%’ , "’" , ‘,’ , ‘/’ , ‘:’ , ‘;’ , ‘<‘ , ‘=’ , ‘>’ , ‘@’ и "`" больше не избегают.

Очистить кэш регулярного выражения.

Исключение подняло, когда строка прошел к одной из функций, вот не действительного регулярного выражения (например, это могло бы содержать непревзойденные круглые скобки), или когда некоторая другая ошибка происходит во время компиляции или соответствия. Если строка не содержит совпадений для шаблона, это никогда не является ошибкой. Ошибка сущность имеет следующие дополнительные атрибуты:

Неформатированное сообщение об ошибке.

Шаблон регулярного выражения.

Индекс в pattern, где сбой компиляции (может быть None ).

Строка, соответствующая pos (может быть None ).

Столбец, соответствующий pos (может быть None ).

Изменено в версии 3.5: Добавлены дополнительные атрибуты.

Объекты регулярного выражения¶

Скомпилированные объекты регулярного выражения поддерживают следующие методы и атрибуты:

Pattern. search ( string [ , pos [ , endpos ] ] ) ¶

Отсканировать string в поисках первого местоположения, в котором это регулярное выражение приводит к совпадению, и возвращает соответствующую объект соответствия . Возвращает None если ни одна позиция в строка не соответствует шаблону; следует отметить, что это отличается от поиска совпадения нулевой длины в какой-то точке строка.

Необязательный второй параметр pos дает индекс в строка, где должен начинаться поиск; по умолчанию используется значение 0 . Это не полностью эквивалентно нарезке строка; символ образца ‘^’ соответствует в реальном начале строка и в позициях сразу после newline, но не обязательно в индексе, где поиск должен начаться.

Необязательный параметр endpos ограничивает степень найденный строка; это будет как если бы строка длиной в endpos символов, так что только символы от pos до endpos — 1 будут найденный для соответствия. Если endpos меньше pos, совпадение не будет найдено; в противном случае, если rx является скомпилированным объектом регулярного выражения, rx.search(string, 0, 50) эквивалентен rx.search(string[:50], 0) .:

Если ноль или больше знаков начинается с string соответствуют этому регулярному выражению, возвращает соответствующий объект соответствия . Возвращает None , если строка не соответствует шаблону; следует отметить, что это отличается от совпадения нулевой длины.

Необязательные параметры pos и endpos имеют то же значение, что и для метода search() :

Если требуется найти совпадение в любом месте string, используйте команду search() (см. также search() против match() ).

Pattern. fullmatch ( string [ , pos [ , endpos ] ] ) ¶

Если целая string соответствует этому регулярному выражению, возвращает соответствующий объект соответствия . Возвращает None , если строка не соответствует шаблону; следует отметить, что это отличается от совпадения нулевой длины.

Необязательные параметры pos и endpos имеют то же значение, что и для метода search() :

Добавлено в версии 3.4.

Идентична функции split() с использованием скомпилированного шаблона.

Pattern. findall ( string [ , pos [ , endpos ] ] ) ¶

Аналогично функции findall() , используя скомпилированный шаблон, но также принимает необязательные параметры pos и endpos, которые ограничивают область поиска, как для search() .

Pattern. finditer ( string [ , pos [ , endpos ] ] ) ¶

Аналогично функции finditer() , используя скомпилированный шаблон, но также принимает необязательные параметры pos и endpos, которые ограничивают область поиска, как для search() .

Идентична функции sub() с использованием скомпилированного шаблона.

Идентичен функции subn() , используя скомпилированный шаблон.

Флаги соответствия regex. Это — комбинация флагов, данных compile() , любым действующим флагам (. ) в образце и неявным флагам, таким как UNICODE , если шаблон — Юникод строка.

Число групп захвата в шаблоне.

Словарь, сопоставляющий любые имена символических групп, определенные с помощью (?P<id>) , с номерами групп. Словарь пуст, если никакие символические группы не были используемый в образце.

Шаблон строка, из которого был скомпилирован объект шаблона.

Изменено в версии 3.7: Добавлена поддержка copy.copy() и copy.deepcopy() . Скомпилированные объекты регулярного выражения считаются атомарными.

Match объекты¶

Match объекты всегда имеют логическую значение True . Начиная с match() и search() возвращает None , когда там не идет ни в какое сравнение, вы можете проверить, был ли соответсвие с простой if инструкцией:

Объекты сопоставления поддерживают следующие методы и атрибуты:

Match. expand ( template ) ¶

Возвращает строку, полученные путем замены обратной косой черты на шаблоне строка template, как это делается методом sub() . Сбеги, такие как \n , преобразуются в соответствующие символы, а числовые обратные ссылки ( \1 , \2 ) и именованные обратные ссылки ( \g<1> , \g<name> ) заменяются содержимым соответствующей группы.

Изменено в версии 3.5: Несопоставленные группы заменяются пустым строка.

Возвращает одну или несколько подгрупп соответствия. Если есть единственный аргумент, результат — единственный строка; при наличии нескольких аргументов результатом является кортеж с одним элементом на аргумент. Без аргументов, дефолтов group1 к нолю (целый матч — возвращенный). Если аргумент groupN — ноль, соответствующий возвращает значение — все соответствие строка; если он находится в инклюзивном диапазоне [1..99], он строка соответствует соответствующей группе, заключенной в скобки. Если номер группы является отрицательным или превышает число групп, определенных в шаблоне, возникает исключение IndexError . Если группа содержится в части массива, которая не соответствует, соответствующий результат является None . Если группа содержится в части шаблона, совпадающей несколько раз, последнее совпадение является возвращенный.

Если регулярное выражение использует синтаксис (?P<name>. ) , аргументы groupN также могут быть строки идентифицирующими группы по имени группы. Если аргумент строка не используемый как название группы в образце, исключение IndexError поднято.

Умеренно сложный пример:

Именованные группы также могут ссылаться по их индексу:

Если группа совпадает несколько раз, доступно только последнее совпадение:

Это идентично m.group(g) . Это упрощает доступ к отдельной группе из совпадения:

Добавлено в версии 3.6.

Возвращает кортеж, содержащий все подгруппы матча, от 1 до, однако многие группы находятся в шаблоне. Аргумент default является используемый для групп, которые не участвовали в матче; по умолчанию используется значение None .

Если сделать десятичный знак и все после него необязательными, не все группы могли бы участвовать в матче. Эти группы будут по умолчанию иметь значение None , если не задан аргумент default:

Возвращает словарь, содержащий все подгруппы named матча, включенного подназванием группы. Аргумент default является используемый для групп, которые не участвовали в матче; по умолчанию используется значение None . Например:

Возвращает индексы начала и конца подстроки, совпадающие по group; group по умолчанию равен нулю (что означает всю совпадающую подстроку). Возвращает -1 если group существует, но не способствовал матчу. Для объекта соответствия m и группы g, которая вносит вклад в соответствие, подстрока, совпадающая по группе g (эквивалентная m.group(g) ), имеет значение:

Обратите внимание, что m.start(group) будет равно m.end(group) , если group соответствует null строка. Например, после m = re.search(‘b(c?)’, ‘cba’) , m.start(0) равно 1, m.end(0) равно 2, m.start(1) и m.end(1) оба 2, и m.start(2) вызывает исключение IndexError .

Пример удаления remove_this из адресов электронной почты:

Для соответствующего m возвращает 2-кортеж (m.start(group), m.end(group)) . Отметим, что если group не способствовал матчу, то это (-1, -1) . group по умолчанию равен нулю, весь матч.

Значение pos, который был передан к search() или методу match() объект регулярного выражения . Это индекс в строка, на котором движок RE начал поиск соответствия.

Значение endpos, который был передан к search() или методу match() объект регулярного выражения . Это индекс в строка, за которым двигатель RE не выйдет.

Целочисленный индекс последней сопоставленной группы захвата или None , если группа не была сопоставлена вообще. Например, выражения (a)b , ((a)(b)) и ((ab)) будут иметь lastindex == 1 , если применяются к строка ‘ab’ , в то время как выражение (a)(b) будет иметь lastindex == 2 , если применяется к тому же строка.

Имя последней сопоставленной группы захвата или None , если у группы не было имени или если группа не была сопоставлена вообще.

Переданная строка match() или search() .

Изменено в версии 3.7: Добавлена поддержка copy.copy() и copy.deepcopy() . Совпадающие объекты считаются атомарными.

Примеры регулярных выражений¶

Проверка на наличие пары¶

В этом примере мы будем использовать следующую вспомогательную функцию, чтобы отображать объекты соответствия немного более грациозно:

Предположим, что вы пишете покерную программу, где рука игрока представлена как 5-символ строка с каждым символ, представляющим карту, «a» для туза, «k» для короля, «q» для королевы, «j» для джека, «t» для 10, и «2» до «9», представляющий карту с этим значение.

Чтобы увидеть, является ли данный строка действительной рукой, можно сделать следующее:

Последняя рука, "727ak" , содержала пару или две одинаковые ценные карты. Чтобы сопоставить это с регулярным выражением, можно использовать обратные ссылки как таковые:

Чтобы узнать, из какой карты состоит пара, можно использовать метод group() объекта соответствия следующим образом:

Имитация scanf()¶

Python в настоящее время не имеет эквивалента scanf() . Регулярные выражения, как правило, более мощные, хотя и более подробные, чем scanf() формат строки. В таблице ниже представлены более или менее эквивалентные сопоставления между маркерами scanf() и регулярными выражениями.

scanf() Токен Регулярное выражение
%c .
%5c .
%d [-+]?\d+
%e , %E , %f , %g [-+]?(\d+(\.\d*)?|\.\d+)([eE][-+]?\d+)?
%i [-+]?(0[xX][\dA-Fa-f]+|0[0-7]*|\d+)
%o [-+]?[0-7]+
%s \S+
%u \d+
%x , %X [-+]?(0[xX])?[\dA-Fa-f]+

Извлечение имени файла и чисел из строка-лайка:

можно использовать формат scanf() , например:

Эквивалентным регулярным выражением будет:

search() против match()¶

Python предлагает две разные примитивные операции, основанные на регулярных выражениях: re.match() проверяет соответствие только в начале строка, в то время как re.search() проверяет соответствие в любой точке строка (это то, что Perl делает по умолчанию).

Регулярные выражения, начинающиеся с ‘^’ , могут быть используемый с search() , чтобы ограничить матч в начале строка:

Обратите внимание однако, что в методе match() MULTILINE только соответствует в начале строка, тогда как использование search() с регулярным выражением, начинающимся с ‘^’ , будет соответствовать в начале каждой линии:

Создание телефонной книги¶

split() разбивает строка на список, разделенный переданным шаблоном. Этот метод неоценим для преобразования текстовых данных в структуры данных, которые могут быть легко считаны и изменены Python, как показано в следующем примере, создающем телефонную книгу.

Во-первых, вот входные данные. Обычно это может прибыть из файла, здесь мы используем трижды указанный синтаксис строка

Записи разделяются одной или несколькими новыми строками. Теперь мы преобразуем строка в список с каждой непустой строкой, имеющей свою собственную запись:

Наконец, разделите каждую запись на список с именем, фамилией, телефонным номером и адресом. Мы используем параметр maxsplit split() , потому что у адреса есть места, наш сильный шаблон, в it:

Шаблон 😕 соответствует двоеточию после фамилии, так что он не появляется в списке результатов. С maxsplit 4 мы могли бы отделить номер дома от названия улицы:

Выравнивание текста¶

sub() заменяет каждое вхождение шаблона на строка или результат функции. Этот пример демонстрирует использование sub() с функцией к «munge» тексту, или рандомизируйте заказ всех знаков в каждом слове предложения за исключением первых и последних знаков:

Найти все наречия¶

findall() совпадает с всеми вхождениями шаблона, а не только с первым, как search() . Например, если писатель хотел найти все наречия в каком-то тексте, они могли бы использовать findall() следующим образом:

Поиск всех наречий и их позиций¶

Если требуется больше информации обо всех совпадениях шаблона, чем сопоставляемого текста, finditer() полезно, так как он предоставляет объекты соответствия вместо строки. Продолжая предыдущий пример, если бы писатель хотел найти все наречия и их позиции в каком-то тексте, они бы использовали finditer() следующим образом:

Необработанная строковая нотация¶

Сырое примечание ( r"text" ) строка сохраняет регулярные выражения нормальными. Без него каждая обратная косая черта ( ‘\’ ) в регулярном выражении должна была бы иметь префикс с другим, чтобы избежать ее. Например, две следующие строки код функционально идентичны:

Когда каждый хочет соответствовать обратной косой черте литерал, ее нужно избежать в регулярном выражении. С сырым примечанием строка это означает r"\\" . Без сырого примечания строка нужно использовать "\\\\" , делая следующие линии код функционально идентичными:

Написание токенизатора¶

Токенизатор или сканер анализирует строку для классификации групп символов. Это полезный первый шаг в написании компилятора или интерпретатор.

Текстовые категории задаются регулярными выражениями. Метод состоит в том, чтобы объединить их в единое основное регулярное выражение и закольцевать последовательные совпадения:

Добавить комментарий

Ваш адрес email не будет опубликован. Обязательные поля помечены *