Перейти к содержимому

Как сделать проверку пароля в python

  • автор:

Check the password strength in Python

Lets us see how to check the strength of the password in Python in this tutorial. Here in this tutorial, we are going to learn how to classify a password according to its strength.

Generally, we think of using functions (with isdigit(), islower(), isupper()), importing ascii_lower, ascii_upper, digits from string, importing packages like PasswordPolicy from password_strength and program accordingly.

Let us now do it in the simplest way…

To classify a password according to its strength

Now, let us do this using regular expressions .

So initially, a password must have 8 characters or more than that.
To have a strong password we must have a digit, a lowercase, an uppercase and a special character or else it is considered as weak .

Regular expression for a strong password

This is the regular expression for strong password.

“(?=.*\d)” says that it must contain one digit similarly we say that for lowercase, uppercase and special characters.
Here ,” ” tells that its length is at least 8 characters and maximum of 30 characters.

R egular expression for weak password

This is the regular expression for weak password.

Here * indicates zero or more than that.

To use regular expressions we need to import re.

Let us have a look on code now.

So, here in our code, we have used the Python len() to know the length of the entered string(input given by the user).
If the length of the string is more than 8 characters then only it is considered as a valid string.

re.match()

The “match()” is a function from module “re”. This function helps to match regular expressions with the string.
Here, in our code, the re.match() function is allowed to return a Boolean value because we are saying the re.match() to return Boolean by putting it in bool.
Generally, the re.match() function returns a match object on success and None on failure.

Now, let us see the output.

OUTPUT:

Here are our three outputs.

From our code, it is understood that we are detecting whether a string contains digits, alphabets, special characters or not.
Here is the link for you to detect if a string contains special characters or not.

So we have learned how to check the password strength in Python with example.

One response to “Check the password strength in Python”

Write a program that checks the strength of a password. The password is strong if it has…
At least 1 letter between [a-z] and 1 letter between [A-Z]
At least 1 number between [0-9]
At least 1 character from [!@#$%^&*]
Minimum length of 6 characters
Print out whether or not the password is strong by looping through each character of the password individually. Loop the whole program until the user inputs a password that is strong enough.

Проверка сложности паролей на Python

Пользователи очень любят простые пароли. Причины этого могут быть разные — кто-то просто не задумывается о сложности пароля, кому-то лень запоминать, а кому-то просто нравится когда в качестве пароля используется распространенное, но крутое слово.

Адекватной реакцией на эту проблему со стороны разработчиков является проверка пользовательских паролей и, соответственно если пароль слишком прост, предложение создать пароль посерьезней. Давайте рассмотрим как можно реализовать наиболее распространенные проверки.

Начальные действия

Создадим свой класс ошибки ValidationError, и все функции валидации будем строить по следующему принципу: если пароль валиден — функция просто молча отрабатывает и не возвращает ничего, если пароль не валиден, то функция будет выбрасывать нашу ошибку валидации. Для удобства я буду запускать проверку функций валидаций используя pytest.

Валидация по формату пароля

Самый простой способ это проверить пароль по регулярному выражению. Наиболее частые требования — проконтролировать минимальную длину пароля, наличие символов в верхнем и нижнем регистрах, наличие в пароле чисел и, иногда, спецсимволов.

Валидация по списку наиболее часто встречающихся паролей

Валидация по регулярным выражением это здорово, но я думаю у вас вызвал некоторое подозрение пароль 5qWerty5, который формально проходит нашу проверку. А ведь кроме qwerty существует еще тысячи подобных слов, которые очень любят использовать в качестве паролей пользователи. password, iloveyou, football. тысячи их. Хорошо бы составить список таких слов и проверять не находится ли присланный нам пароль среди них. Хорошая новость — есть на свете такой замечательный человек по имени Royce Williams, который уже собрал тысячи таких паролей. Весь список доступен на gist.

Мы можем скачать архив, который содержит текстовый файл с паролями в следующем формате frequency:sha1-hash:plain, то есть — частота встречаемости пароля, его хеш, и собственно сам пароль как он есть. Давайте напишем функцию, которая будет открывать файл со списком и, итерируясь по строкам, сверять наш пароль с очередным паролем в списке:

Что ж наша функция легко находит такие очевидные слова типа qwerty, но что если пользователь будет не так прост, и на наше замечание, что его пароль слишком очевиден, скажем, просто добавит куда нибудь точку или поставит пару цифр вначале и в конце: (вставить результаты тестов?) qwert.y, 0qwerty0 или даже q.w.e.r.t.y.?

Добавим эти проверки в тест:

Такие очевидные хаки наш валидатор уже не в состоянии отловить.

В качестве решения можно было бы, конечно же, попробовать вставить удаление из строки точек (и/или других спецсимволов), например как-то так password = password.replace(‘.’, ») . Однако всем понятно что такой путь, мягко говоря, не очень эстетичный и правильный. Вместо этого можно воспользоваться модулем стандартной библиотеки python difflib.

Как следует из описания — этот модуль предоставляет классы и функции для сравнения последовательностей, что нам отлично подходит — ведь строки в python обладают свойствами последовательностей. Давайте рассмотри поближе объект difflib.SequenceMatcher .

Класс SequenceMatcher принимает на вход две последовательности и предоставляет несколько методов для оценки их сходства. Нас интересует метод ratio() который возвращает число в диапазоне [0,1] характеризующее «похожесть» двух последовательностей, где 1 соответствует двум абсолютно одинаковым последовательностям, а 0 абсолютно разным.

Перепишем нашу функцию валидации следующим образом:

max_similarity — характеризует максимально допустимое сходство, увлекаться и слишком занижать этот параметр не стоит, иначе ваш валидатор будет улавливать малейшие совпадения вплоть до пары символов. По моему опыту, значение 0.7 это минимальный порог ниже которого опускаться не стоит, при этом порог 0.75 уже пропустит вот такой пароль ‘q.w.e.r.t.y’ , так что определите размер этого параметра для себя сами.

Кроме того, здесь я использую функцию:

dropwhile(lambda x: x.startswith(‘#’), f)

из модуля itertools для того, чтобы пропустить закомментированные строки вначале файла common-passwords.txt, впрочем их можно было просто удалить вручную.

Протестируем наш переписанный валидатор:

Валидация по использованию в качестве пароля других полей

Итак, мы обозначили необходимый формат пароля и проверили его, чтобы он не был слишком очевидным. Другим распространенным случаем является использование в качестве пароля значения другого атрибута пользователя. Например, если пользователь просто скопирует в поле пароля свой email или логин. Для определения таких случаев можно воспользоваться тем же способом, что мы использовали для определения похожих паролей — объектом d ifflib.SequenceMatcher , только в этот раз мы будем сравнивать пароль со значением других полей:

Здесь мы разделяем пароль на части на части по шаблону \W+ , под который подходят все нестандартные символы (то есть не включающие в себя буквы, цифры и нижнее подчеркивание), для случаев, когда пользователь может использовать в качестве пароля часть своего имейла без домена. Например при использовании в качестве пароля имейла someemailname@gmail.com получим следующие части: [‘someemailname’, ‘gmail’, ‘com’, ‘someemailname@gmail.com’].

Проверим как работает наша функция:

Перечисленных способов, как правило, будет достаточно, впрочем, и ими тоже не стоит слишком увлекаться, иначе дружелюбность вашего интерфейса для пользователей станет примерно вот такой:

Так что, выбор какие способы и с какой степенью сложности использовать в своих проектах зависит лишь от серьезности вашего проекта, необходимой степени защиты. ну и пожалуй от ваших садистских наклонностей.

Tutorial

You test your passwords using the Policy object that controls what kind of password is acceptable in your system.

First, create the Policy object and define the rules that apply to passwords in your system:

Now, when you have the PasswordPolicy object, you can use it to test your passwords, and it will tell you which tests have failed:

This tells us that 2 tests have failed: password is not long enough, and it does not have enough special characters. You can use this information to tell the user what precisely is wrong with their password.

Empty list tells us that this password is alright.

This test, however, enabled uses to use passwords that have a lot of repetition.

So-Called Entropy Bits

Here’s a test that’s even better. You don’t really need to define complex rules with special characters and stuff. All you actually need is a password that’s long enough, complex enough, and easy to remember (see xkcd and Article: Everything We’ve Been Told About Passwords Is Wrong).

So, instead of defining all these rules, let’s just require the password to be complex enough. Entropy bits is something that defines how much variety does your password have. ‘01111010010011’ is long enough, but has only 2 entropy bits: that’s how many bits you need to store its alphabet. However, a password that uses plenty of characters has more entropy.

This password is not long enough, or secure enough, but has enough entropy: its vocabulary has 10 different characters. Put this test together with other requirements to make sure there’s no repetition in your passwords.

Complexity

Entropy bits are important, but difficult to understand. An even better, more intuitive test, is to require the password to be «complex enough». Complexity is a number in the range of 0.00..0.99. Good, strong passwords start at 0.66.

Let’s first see how different passwords score:

So, 0.66 will be a very good indication of a good password. Let’s implement our policy:

One good thing about using strength is that it allows users to use national aplhabets with passwords, which are most secure:

Notice how in the last example we use a different approach: policy.password() analyzes the password, and then we can both get its .strength() , and .test() it according to the current policy.

PasswordPolicy

Perform tests on a password.

Init Policy

Init password policy with a list of tests

Init password policy from a dictionary of test definitions.

A test definition is simply:

Test name is just a lowercased class name.

Bundled Tests

These objects perform individual tests on a password, and report True of False .

tests.EntropyBits(bits)

Test whether the password has >= bits entropy bits.

Entropy bits is the number of bits that is required to store the alphabet that’s used in a password. It’s a measure of how long is the alphabet.

tests.Length(length)

Tests whether password length >= length

tests.NonLetters(count)

Test whether the password has >= count non-letter characters

tests.NonLettersLc(count)

Test whether the password has >= count non-lowercase characters

tests.Numbers(count)

Test whether the password has >= count numeric characters

tests.Special(count)

Test whether the password has >= count special characters

tests.Strength(strength, weak_bits=30)

Test whether the password has >= strength strength.

A password is evaluated to the strength of 0.333 when it has weak_bits entropy bits, which is considered to be a weak password. Strong passwords start at 0.666.

tests.Uppercase(count)

Test whether the password has >= count uppercase characters

Testing

After the PasswordPolicy is initialized, there are two methods to test:

PasswordPolicy.password

Get password stats bound to the tests declared in this policy.

If in addition to tests you need to get statistics (e.g. strength) — use this object to double calculations.

See PasswordStats for more details.

PasswordPolicy.test

Perform tests on a password.

Shortcut for: PasswordPolicy.password(password).test() .

Custom Tests

ATest is a base class for password tests.

To create a custom test, just subclass it and implement the following methods:

  • init() that takes configuration arguments
  • test(ps) that tests a password, where ps is a PasswordStats object.

PasswordStats

PasswordStats allows to calculate statistics on a password.

It considers a password as a unicode string, and all statistics are unicode-based.

PasswordStats.alphabet

Get alphabet: set of used characters

PasswordStats.alphabet_cardinality

Get alphabet cardinality: alphabet length

PasswordStats.char_categories

Character count per top-level category

The following top-level categories are defined:

  • L: letter
  • M: Mark
  • N: Number
  • P: Punctuation
  • S: Symbol
  • Z: Separator
  • C: Other
PasswordStats.char_categories_detailed

Character count per unicode category, detailed format.

PasswordStats.combinations

The number of possible combinations with the current alphabet

PasswordStats.count(*categories)

Count characters of the specified classes only

PasswordStats.count_except(*categories)

Count characters of all classes except the specified ones

PasswordStats.entropy_bits

Get information entropy bits: log2 of the number of possible passwords

PasswordStats.entropy_density

Get information entropy density factor, ranged <0 .. 1>.

This is ratio of entropy_bits() to max bits a password of this length could have. E.g. if all characters are unique — then it’s 1.0. If half of the characters are reused once — then it’s 0.5.

PasswordStats.length

Get password length

PasswordStats.letters

Count all letters

PasswordStats.letters_lowercase

Count lowercase letters

PasswordStats.letters_uppercase

Count uppercase letters

PasswordStats.numbers
PasswordStats.repeated_patterns_length

Detect and return the length of repeated patterns.

You will probably be comparing it with the length of the password itself and ban if it’s longer than 10%

PasswordStats.sequences_length

Detect and return the length of used sequences:

  • Alphabet letters: abcd.
  • Keyboard letters: qwerty, etc
  • Keyboard special characters in the top row:

Intro to Regexes & Strong Password Detection in Python

Akeel Ahamed

Regular Expressions (or Regexes) are huge time-savers, not just for software users but also for Programmers and Data Scientists. Tech writer Cory Doctorow argues that even before learning to program, we should be learning regular expressions:

“Knowing [regular expressions] can mean the difference between solving a problem in three steps and solving it in 3,000 steps. When you’re a nerd, you forget that the problems you solve with a couple keystrokes can take other people days of tedious, error-prone work to slog through.”

Regular Expressions are a mini language for specifying text patterns. In this article I will give you a brief introduction to using regular expressions in Python and apply it to a real world problem of creating a password detector.

To demonstrate how much time we can save by having basic knowledge of regexes, let me give you an example of the code used to extract phone numbers in a string of text with and without using regular expressions. For simplicity, let us assume that a phone is valid if it is in the format ddd-ddd-dddd.

The above code shows how much effort is needed to simply print all the phone numbers in a string. Now let us see how regexes can simplify the problem:

Fantastic! In just 3 lines of code, we managed to do the same task!

Hopefully, now that I have your attention, let us explore how we can use regular expressions in Python.

Regular Expressions in Python

Before we move further, I would like to explain the difference between raw strings and normal strings.

A raw string is specified using ‘r’ before beginning the string in Python. It treats the backslash (\) as a literal character. Recall that escape characters in Python use the backslash (\).The string value ‘\n’ represents a single newline character, not a backslash followed by a lowercase n.

In order to specify a backslash followed by a lowercase n, you need to enter the escape character ‘\\’ to print a single backslash, followed by ‘n’. Thus, ‘\\n’ is the string that represents a backslash followed by a lowercase ‘n’. However, by inserting an ‘r’ before the first quote of the string, you can mark the string as a raw string, which does not escape characters.

Python has a built in module to work with regular expressions called re. Let us now explore some of the commonly used methods in the re module:

  • re.match(pattern, string): matches a pattern specified at the beginning of a string and returns a match object if the pattern is present. Otherwise, it returns ‘None’.

Since the output of re.match is a match object, we can use the group() method to return the matched expressions.

Optionally, you can also specify a pattern separately using the re.compile(pattern) function that takes the pattern as an argument.

  • re.search(pattern, string): matches only the first occurence of a pattern in the string.
  • re.findall(pattern, string): returns all occurrences of the searched pattern within a string (in a list format).

This is more powerful than the latter two methods and is one that I prefer using.

  • re.sub(pattern, replacement, string):returns the string obtained by replacing the occurrences of pattern in the string, by replacement. If there is no pattern found, the original string is returned.

Metacharacters & Special Sequences

Regular expressions in general can be specified using a combination of metacharacters and special sequences.

Square brackets [abc]

Square brackets match any character between the brackets (such as a, b or c) in a string.

In the three examples above, string 1 returns two matches, string 2 returns four matches and the last string has no match (as there are no letters a,b or c in string 3).

  • Metacharacters such as “[] . ^ $ + ? <> () \ | .” lose their meaning inside square brackets. For example, [(*+)] will match any instance of the literal characters ‘[’, ‘(’, ‘*’, ‘+’, ‘)’ or ‘]’.

Period (.)

The period matches any one character, except the newline (\n) character (similar to a wildcard character).

Caret (^)

The caret symbol checks if a string begins with a certain character.

Dollar($)

The dollar symbol checks if a string ends with a certain character.

Overtime, between the caret and dollar symbols, it is easy to forget which one comes first. A mnemonic I found to be useful is “Carrots cost dollars!”.

The start symbol matches zero or more occurences of the pattern to the left of it.

The plus symbol matches one or more occurences of the pattern to the left of it.

Braces <>

Taking the pattern r this will match at least x and at most y repetitions of the pattern ‘r’.

Alternation |

The alternation or the “or” operator. Suppose A and B are regular expressions, then A|B will match instances that contain either the expression A or B.

Grouping ()

The parantheses () group together the expression contained inside them. For example, the expression (a|b|c)xy will match all instances containing the characters “a” or “b” or “c”, followed by “xy”.

Question Mark ?

The question mark symbol matches zero or one occurences of the pattern to the left of it.

Backslash \

A backslash is used to escape various characters including all the metacharacters discussed. By ‘escape’, we mean it will invoke the next character called. For example, \$a will match instances with a “$” followed by “a” and is not interpreted by the regex engine in a special way.

However, there are some special sequences that the regex engine does interpret in a special way. Special Sequences make commonly used patterns much easier to write. Some of them are:

Matches any numeric digit from 0 to 9; this is equivalent to the class [0-9] .

Matches any character that is not a numeric digit from 0 to 9; this is equivalent to the class [^0-9] .

Matches any space, tab or newline character (i.e. whitespace characters); this is equivalent to the class [ \t\n\r\f\v] .

Matches any non-whitespace character; this is equivalent to the class [^ \t\n\r\f\v] .

Matches any letter, numeric digit or underscore character (i.e. alphanumeric characters); this is equivalent to the class [a-zA-Z0-9_] .

Matches any non-alphanumeric character; this is equivalent to the class [^a-zA-Z0-9_] .

A more detailed list of special characters can be seen here.

  • Special Sequences are accepted inside the square brackets metacharacter. For example, [\w] will match any instance of a “word” character (letter, numeric digit or underscore character).

Review Exercise

Let us now use regular expressions to extract a grocery list from a string of text. Before looking at the method I have used, I would encourage you to try this out first using the string given:

Password Detector Program

Using our knowledge of regexes, we will now build a small program to detect the strength of a password entered.

The password strength is based on 4 main checks (feel free to add more if you like):

  1. The password should have at least 8 characters.
  2. The password should contain at least one lowercase character.
  3. The password should contain at least one uppercase character.
  4. The password should contain at least one digit.

Summary Notes

There has been a lot covered in this article so far. Here is a brief review of the steps to use regular expressions in Python:

  1. Import the regex module with import re .
  2. Create a Regex object with the re.compile() function. (Remember to use a raw string.)
  3. Pass the string you want to search into the Regex object’s search() method. This returns a Match object.
  4. Call the Match object’s group() method to return a string of the actual matched text.

Notice that “.group()” method did not return all the matches of phone numbers in the string. To return all the matched text corresponding to the pattern searched, we can use the “findall()” method:

Note that you can also choose to skip steps 2,3 and 4 and instead use the findall() method directly.

The following is a great review of the symbols learned in this article extracted from here:

  • The ? matches zero or one of the preceding group.
  • The * matches zero or more of the preceding group.
  • The + matches one or more of the preceding group.
  • The matches exactly n of the preceding group.
  • The matches n or more of the preceding group.
  • The <,m>matches 0 to m of the preceding group.
  • The matches at least n and at most m of the preceding group.
  • ? or *? or +? performs a nongreedy match of the preceding group.
  • ^spam means the string must begin with spam.
  • spam$ means the string must end with spam.
  • The . matches any character, except newline characters.
  • \d , \w , and \s match a digit, word, or space character, respectively.
  • \D , \W , and \S match anything except a digit, word, or space character, respectively.
  • [abc] matches any character between the brackets (such as a, b, or c).
  • [^abc] matches any character that isn’t between the brackets.

Regex Practice

If you would like to practice the concepts learned in this article, some helpful resources are regex101 and regexone.

The Jupyter Notebook for this article can be found here.

Hope you enjoyed learning about regular expressions and have a fabulous day!

Добавить комментарий

Ваш адрес email не будет опубликован. Обязательные поля помечены *