Перейти к содержимому

Как изменить кодировку csv файла на utf 8 python

  • автор:

csv — CSV File Reading and Writing¶

The so-called CSV (Comma Separated Values) format is the most common import and export format for spreadsheets and databases. CSV format was used for many years prior to attempts to describe the format in a standardized way in RFC 4180. The lack of a well-defined standard means that subtle differences often exist in the data produced and consumed by different applications. These differences can make it annoying to process CSV files from multiple sources. Still, while the delimiters and quoting characters vary, the overall format is similar enough that it is possible to write a single module which can efficiently manipulate such data, hiding the details of reading and writing the data from the programmer.

The csv module implements classes to read and write tabular data in CSV format. It allows programmers to say, “write this data in the format preferred by Excel,” or “read data from this file which was generated by Excel,” without knowing the precise details of the CSV format used by Excel. Programmers can also describe the CSV formats understood by other applications or define their own special-purpose CSV formats.

The csv module’s reader and writer objects read and write sequences. Programmers can also read and write data in dictionary form using the DictReader and DictWriter classes.

The Python Enhancement Proposal which proposed this addition to Python.

Module Contents¶

The csv module defines the following functions:

csv. reader ( csvfile , dialect = ‘excel’ , ** fmtparams ) ¶

Return a reader object which will iterate over lines in the given csvfile. csvfile can be any object which supports the iterator protocol and returns a string each time its __next__() method is called — file objects and list objects are both suitable. If csvfile is a file object, it should be opened with newline=» . 1 An optional dialect parameter can be given which is used to define a set of parameters specific to a particular CSV dialect. It may be an instance of a subclass of the Dialect class or one of the strings returned by the list_dialects() function. The other optional fmtparams keyword arguments can be given to override individual formatting parameters in the current dialect. For full details about the dialect and formatting parameters, see section Dialects and Formatting Parameters .

Each row read from the csv file is returned as a list of strings. No automatic data type conversion is performed unless the QUOTE_NONNUMERIC format option is specified (in which case unquoted fields are transformed into floats).

A short usage example:

Return a writer object responsible for converting the user’s data into delimited strings on the given file-like object. csvfile can be any object with a write() method. If csvfile is a file object, it should be opened with newline=» 1. An optional dialect parameter can be given which is used to define a set of parameters specific to a particular CSV dialect. It may be an instance of a subclass of the Dialect class or one of the strings returned by the list_dialects() function. The other optional fmtparams keyword arguments can be given to override individual formatting parameters in the current dialect. For full details about dialects and formatting parameters, see the Dialects and Formatting Parameters section. To make it as easy as possible to interface with modules which implement the DB API, the value None is written as the empty string. While this isn’t a reversible transformation, it makes it easier to dump SQL NULL data values to CSV files without preprocessing the data returned from a cursor.fetch* call. All other non-string data are stringified with str() before being written.

A short usage example:

Associate dialect with name. name must be a string. The dialect can be specified either by passing a sub-class of Dialect , or by fmtparams keyword arguments, or both, with keyword arguments overriding parameters of the dialect. For full details about dialects and formatting parameters, see section Dialects and Formatting Parameters .

csv. unregister_dialect ( name ) ¶

Delete the dialect associated with name from the dialect registry. An Error is raised if name is not a registered dialect name.

csv. get_dialect ( name ) ¶

Return the dialect associated with name. An Error is raised if name is not a registered dialect name. This function returns an immutable Dialect .

Return the names of all registered dialects.

csv. field_size_limit ( [ new_limit ] ) ¶

Returns the current maximum field size allowed by the parser. If new_limit is given, this becomes the new limit.

The csv module defines the following classes:

class csv. DictReader ( f , fieldnames = None , restkey = None , restval = None , dialect = ‘excel’ , * args , ** kwds ) ¶

Create an object that operates like a regular reader but maps the information in each row to a dict whose keys are given by the optional fieldnames parameter.

The fieldnames parameter is a sequence . If fieldnames is omitted, the values in the first row of file f will be used as the fieldnames. Regardless of how the fieldnames are determined, the dictionary preserves their original ordering.

If a row has more fields than fieldnames, the remaining data is put in a list and stored with the fieldname specified by restkey (which defaults to None ). If a non-blank row has fewer fields than fieldnames, the missing values are filled-in with the value of restval (which defaults to None ).

All other optional or keyword arguments are passed to the underlying reader instance.

Changed in version 3.6: Returned rows are now of type OrderedDict .

Changed in version 3.8: Returned rows are now of type dict .

A short usage example:

Create an object which operates like a regular writer but maps dictionaries onto output rows. The fieldnames parameter is a sequence of keys that identify the order in which values in the dictionary passed to the writerow() method are written to file f. The optional restval parameter specifies the value to be written if the dictionary is missing a key in fieldnames. If the dictionary passed to the writerow() method contains a key not found in fieldnames, the optional extrasaction parameter indicates what action to take. If it is set to ‘raise’ , the default value, a ValueError is raised. If it is set to ‘ignore’ , extra values in the dictionary are ignored. Any other optional or keyword arguments are passed to the underlying writer instance.

Note that unlike the DictReader class, the fieldnames parameter of the DictWriter class is not optional.

A short usage example:

The Dialect class is a container class whose attributes contain information for how to handle doublequotes, whitespace, delimiters, etc. Due to the lack of a strict CSV specification, different applications produce subtly different CSV data. Dialect instances define how reader and writer instances behave.

All available Dialect names are returned by list_dialects() , and they can be registered with specific reader and writer classes through their initializer ( __init__ ) functions like this:

The excel class defines the usual properties of an Excel-generated CSV file. It is registered with the dialect name ‘excel’ .

class csv. excel_tab ¶

The excel_tab class defines the usual properties of an Excel-generated TAB-delimited file. It is registered with the dialect name ‘excel-tab’ .

class csv. unix_dialect ¶

The unix_dialect class defines the usual properties of a CSV file generated on UNIX systems, i.e. using ‘\n’ as line terminator and quoting all fields. It is registered with the dialect name ‘unix’ .

New in version 3.2.

The Sniffer class is used to deduce the format of a CSV file.

The Sniffer class provides two methods:

sniff ( sample , delimiters = None ) ¶

Analyze the given sample and return a Dialect subclass reflecting the parameters found. If the optional delimiters parameter is given, it is interpreted as a string containing possible valid delimiter characters.

Analyze the sample text (presumed to be in CSV format) and return True if the first row appears to be a series of column headers. Inspecting each column, one of two key criteria will be considered to estimate if the sample contains a header:

  • the second through n-th rows contain numeric values

  • the second through n-th rows contain strings where at least one value’s length differs from that of the putative header of that column.

Twenty rows after the first row are sampled; if more than half of columns + rows meet the criteria, True is returned.

This method is a rough heuristic and may produce both false positives and negatives.

An example for Sniffer use:

The csv module defines the following constants:

Instructs writer objects to quote all fields.

Instructs writer objects to only quote those fields which contain special characters such as delimiter, quotechar or any of the characters in lineterminator.

Instructs writer objects to quote all non-numeric fields.

Instructs the reader to convert all non-quoted fields to type float.

Instructs writer objects to never quote fields. When the current delimiter occurs in output data it is preceded by the current escapechar character. If escapechar is not set, the writer will raise Error if any characters that require escaping are encountered.

Instructs reader to perform no special processing of quote characters.

The csv module defines the following exception:

exception csv. Error ¶

Raised by any of the functions when an error is detected.

Dialects and Formatting Parameters¶

To make it easier to specify the format of input and output records, specific formatting parameters are grouped together into dialects. A dialect is a subclass of the Dialect class having a set of specific methods and a single validate() method. When creating reader or writer objects, the programmer can specify a string or a subclass of the Dialect class as the dialect parameter. In addition to, or instead of, the dialect parameter, the programmer can also specify individual formatting parameters, which have the same names as the attributes defined below for the Dialect class.

Dialects support the following attributes:

A one-character string used to separate fields. It defaults to ‘,’ .

Controls how instances of quotechar appearing inside a field should themselves be quoted. When True , the character is doubled. When False , the escapechar is used as a prefix to the quotechar. It defaults to True .

On output, if doublequote is False and no escapechar is set, Error is raised if a quotechar is found in a field.

A one-character string used by the writer to escape the delimiter if quoting is set to QUOTE_NONE and the quotechar if doublequote is False . On reading, the escapechar removes any special meaning from the following character. It defaults to None , which disables escaping.

Changed in version 3.11: An empty escapechar is not allowed.

The string used to terminate lines produced by the writer . It defaults to ‘\r\n’ .

The reader is hard-coded to recognise either ‘\r’ or ‘\n’ as end-of-line, and ignores lineterminator. This behavior may change in the future.

A one-character string used to quote fields containing special characters, such as the delimiter or quotechar, or which contain new-line characters. It defaults to ‘"’ .

Changed in version 3.11: An empty quotechar is not allowed.

Controls when quotes should be generated by the writer and recognised by the reader. It can take on any of the QUOTE_* constants (see section Module Contents ) and defaults to QUOTE_MINIMAL .

When True , spaces immediately following the delimiter are ignored. The default is False .

When True , raise exception Error on bad CSV input. The default is False .

Reader Objects¶

Reader objects ( DictReader instances and objects returned by the reader() function) have the following public methods:

Return the next row of the reader’s iterable object as a list (if the object was returned from reader() ) or a dict (if it is a DictReader instance), parsed according to the current Dialect . Usually you should call this as next(reader) .

Reader objects have the following public attributes:

A read-only description of the dialect in use by the parser.

The number of lines read from the source iterator. This is not the same as the number of records returned, as records can span multiple lines.

DictReader objects have the following public attribute:

If not passed as a parameter when creating the object, this attribute is initialized upon first access or when the first record is read from the file.

Writer Objects¶

Writer objects ( DictWriter instances and objects returned by the writer() function) have the following public methods. A row must be an iterable of strings or numbers for Writer objects and a dictionary mapping fieldnames to strings or numbers (by passing them through str() first) for DictWriter objects. Note that complex numbers are written out surrounded by parens. This may cause some problems for other programs which read CSV files (assuming they support complex numbers at all).

csvwriter. writerow ( row ) ¶

Write the row parameter to the writer’s file object, formatted according to the current Dialect . Return the return value of the call to the write method of the underlying file object.

Changed in version 3.5: Added support of arbitrary iterables.

Write all elements in rows (an iterable of row objects as described above) to the writer’s file object, formatted according to the current dialect.

Writer objects have the following public attribute:

A read-only description of the dialect in use by the writer.

DictWriter objects have the following public method:

Write a row with the field names (as specified in the constructor) to the writer’s file object, formatted according to the current dialect. Return the return value of the csvwriter.writerow() call used internally.

New in version 3.2.

Changed in version 3.8: writeheader() now also returns the value returned by the csvwriter.writerow() method it uses internally.

Examples¶

The simplest example of reading a CSV file:

Reading a file with an alternate format:

The corresponding simplest possible writing example is:

Since open() is used to open a CSV file for reading, the file will by default be decoded into unicode using the system default encoding (see locale.getencoding() ). To decode a file using a different encoding, use the encoding argument of open:

The same applies to writing in something other than the system default encoding: specify the encoding argument when opening the output file.

Registering a new dialect:

A slightly more advanced use of the reader — catching and reporting errors:

And while the module doesn’t directly support parsing strings, it can easily be done:

If newline=» is not specified, newlines embedded inside quoted fields will not be interpreted correctly, and on platforms that use \r\n linendings on write an extra \r will be added. It should always be safe to specify newline=» , since the csv module does its own ( universal ) newline handling.

Convert CSV to UTF-8 in Python

I am trying to create a duplicate CSV without a header. When I attempt this I get the following error:

I’ve read the python CSV documentation on Unicode and UTF-8 encoding and have implemented it. However, my output file is being generated with no data in it. Not sure what I am doing wrong here.

3 Answers 3

The solution was to simply include two additional parameters to the

The two parameters are encoding =’UTF-8′ and errors=’ignore’. This allowed me to create a duplicate of original CSV without the headers and without the UnicodeDecodeError. Below is the completed code.

isn’t indented, it is out of the scope of the with command, and when it called, then infile and outfile are both closed.

The files should be opened when they are used, not when the functions are defined, so have:

Как изменить кодировку csv файла на utf 8 python

Several errors can arise when an attempt to decode a byte string from a certain coding scheme is made. The reason is the inability of some encoding schemes to represent all code points. One of the most common errors during these conversions is UnicodeDecode Error which occurs when decoding a byte string by an incorrect coding scheme. This article will teach you how to resolve a UnicodeDecodeError for a CSV file in Python.

Why does the UnicodeDecodeError error arise?

The error occurs when an attempt to represent code points outside the range of the coding is made. To solve the issue, the byte string should be decoded using the same coding scheme in which it was encoded. i.e., The encoding scheme should be the same when the string is encoded and decoded.

For demonstration, the same error would be reproduced and then fixed. In the below code, firstly the character a (byte string) is decoded using ASCII encoding successfully. Then an attempt to decode the byte string a\xf1 is made, which led to an error. This is because the ASCII encoding standard only allows representation of the characters within the range 0 to 127. Any attempt to address a character outside this range would lead to the ordinal not-in-range error.

csv — Чтение и запись CSV файлов¶

Так называемый формат CSV (значения, разделённые запятыми) является наиболее распространенным форматом импорта и экспорта электронных таблиц и баз данных. Формат CSV использовался в течение многих лет до попыток описать формат стандартизированным образом в RFC 4180. Отсутствие четко определенного стандарта означает, что в данных, производимых и потребляемых различными приложениями, часто существуют незначительные различия. Эти различия могут вызвать раздражение при обработке файлов CSV из нескольких источников. Тем не менее, хотя разделители и символы кавычек различаются, общий формат достаточно похож, чтобы можно было написать один модуль, который может эффективно манипулировать такими данными, скрывая детали чтения и записи данных от программиста.

Модуль csv реализует классы для чтения и записи табулированных данных в формате CSV. Он позволяет программистам говорить, «запиши данные в формате, предпочитаемом Excel», или «прочти данные из файла, который был создан Excel», не зная точных деталей формата CSV используемого в Excel. Программисты также могут описать CSV форматы, понятные другим приложениям или определить собственные специальные форматы CSV.

Объекты reader и writer модуля csv читают и пишут последовательности. Программисты могут также читать и записывать данные в словарной форме, используя классы DictReader и DictWriter .

PEP 305 — CSV файловый API Предложение по улучшению Python, предложившее это дополнение к Python.

Содержание модуля¶

Модуль csv определяет следующие функции:

csv. reader ( csvfile, dialect=’excel’, **fmtparams ) ¶

Возвращает объект reader, который будет итерироваться по строкам в предоставленном csvfile. csvfile может быть любым объектом, который поддерживает протокол итератор и возвращает строку каждый раз, когда вызывается метод __next__() файлового объекта или объекта списка. Если csvfile является файловым объектом, то его нужно открыть с параметром newline=» . [1] Дополнительный параметр dialect используется для определения ряда параметров, характерных для специфического CSV диалекта. Это может быть сущность подкласса класса Dialect или одной из строк, возвращаемой функцией list_dialects() . Также могут передаваться другие дополнительные ключевые аргументы fmtparams для переопределения отдельных параметров форматирования в текущем диалекте. Подробные сведения о параметрах диалекта и форматировании см. в разделе Диалекты и параметры форматирования .

Каждая строка, считанная из файла csv, возвращается в виде списка строк. Автоматическое преобразование типов данных не выполняется, если не указан параметр формата QUOTE_NONNUMERIC (в этом случае поля без кавычек преобразуются в числа с плавающей точкой).

Короткий пример использования:

Возвращает объект writer, отвечающий за преобразование пользовательских данных с отдельными строками в предоставленный файлоподобный объект. csvfile может быть любым объектом с методом write() . Если csvfile является файловым объектом, его следует открыть с помощью newline=» [1]. Дополнительный параметр dialect используется для определения ряда параметров, характерных для специфического CSV диалекта. Это может быть сущность подкласса класса Dialect или одной из строк, возвращаемой функцией list_dialects() . Также могут передаваться другие дополнительные ключевые аргументы fmtparams для переопределения отдельных параметров форматирования в текущем диалекте. Подробные сведения о параметрах диалекта и форматировании см. в разделе Диалекты и параметры форматирования . Для максимально простого взаимодействия с модулями, реализующими DB API, значение None записывается как пустая строка. В то время как это необратимое преобразование, но оно помогает сделать проще SQL дамп с NULL данными в CSV файлы без предварительной обработки данных, возвращаемых вызовом cursor.fetch* . Все другие нестроковые данные перед записью стрингифицируются функцией str() .

Короткий пример использования:

Связывание dialect с name. name должен быть строкой. dialect может быть определен передав подкласс Dialect или ключевые аргументы fmtparams или оба, с ключевыми аргументами переопределяющими параметры диалекта. Для получения всех подробностей о диалекте и параметрах форматирования, см. раздел Диалекты и параметры форматирования .

csv. unregister_dialect ( name ) ¶

Удаляет диалект, связанный с name зарегистрированного диалекта. Вызывается Error , если name не зарегистрированное имя диалекта.

csv. get_dialect ( name ) ¶

Возвращает диалект, связанный с name. Вызывается Error , если name не зарегистрированное имя диалекта. Функция возвращает неизменяемый Dialect .

Возвращает названия всех зарегистрированных диалектов.

csv. field_size_limit ( [ new_limit ] ) ¶

Возвращает текущий максимальный размер поля, разрешённого парсером. Если new_limit передан, то он становится новым ограничением.

Модуль csv определяет следующие классы:

Создает объект, работающий как обычный reader, но отображающий информацию каждой строки в dict , ключи которого задаются необязательным параметром fieldnames.

Параметр fieldnames является последовательностью . Если fieldnames пропущен, значения в первой строке файла f будут использоваться в качестве имен полей. Независимо от того, как определяются имена полей, словарь сохраняет их первоначальный порядок.

Если у строки больше полей, чем fieldnames, остающиеся данные помещаются в список и хранятся с fieldname, определенным restkey (который по умолчанию None ). Если у не пустой строки меньше полей, чем fieldnames, отсутствующие значения заполняются значениями restval (которые по умолчанию None ).

Все другие дополнительные или ключевые аргументы передаются основной сущности reader .

Изменено в версии 3.6: Возвращаемые строки теперь имеют тип OrderedDict .

Изменено в версии 3.8: Возвращенные строки теперь имеют тип dict .

Короткий пример использования:

Создаёт объект, который работает как обычный writer, но отображает словари на выходные строки. Параметр fieldnames — это последовательность ключей, идентифицирующих порядок, в котором значения словаря, переданные методу writerow() , записываются в файл f. Необязательный параметр restval указывает значение для записи, если в словаре отсутствует ключ в fieldnames. Если словарь, переданный методу writerow() , содержит ключ, не найденный в fieldnames, необязательный параметр extrasaction указывает, какое действие необходимо предпринять. Если установлено значение ‘raise’ (значение по умолчанию), вызывается исключение ValueError . Если установлено значение ‘ignore’ , дополнительные значения в словаре игнорируются. Любые другие необязательные или ключевые аргументы передаются в базовую сущность writer .

Обратите внимание, что в отличие от класса DictReader , параметр fieldnames класса DictWriter не опциональный.

Короткий пример использования:

Класс Dialect является контейнером класса, основанным главным образом на его атрибутах, которые используются для определения параметров для сущности reader или writer .

Класс excel определяет обычные свойства создаваемого Excel файла CSV. Зарегистрирован с именем диалекта ‘excel’ .

class csv. excel_tab ¶

Класс excel_tab определяет обычные свойства произведенного Excel разделённого TAB’ами файла. Он зарегистрирован с именем диалекта ‘excel-tab’ .

class csv. unix_dialect ¶

Класс unix_dialect определяет обычные свойства файла CSV, генерируемого в системах UNIX, т.е. используя ‘\n’ в качестве терминатора строки и заковычивая все поля. Зарегистрирован с диалектным именуемым ‘unix’ .

Добавлено в версии 3.2.

Класс Sniffer используется для вывода формата файла CSV.

В классе Sniffer предусмотрены два метода:

sniff ( sample, delimiters=None ) ¶

Проанализировать данный sample и возвратить подкласс Dialect , отражающий найденные параметры. Если дополнительный параметр предоставлен delimiters, то он интерпретируется как строка, содержащий возможные допустимые символы разделителя.

Анализирует типовой текст (предполагается, что он в формате CSV) и возвращает True , если первая строка, кажется, строкой заголовков колонки.

Модуль csv определяет следующие константы:

Предписывает writer объектам закавычивать все поля.

Предписывает writer объектам закавычивать только те поля, которые содержат специальные символы, такие как delimiter, quotechar или любой из символов в lineterminator.

Предписывает writer объектам закавычивать все нечисловые поля.

Дает указание reader преобразовывать все незаковыченные поля в тип float.

Предписывает writer объектам никогда не закавычивать поля. Когда текущий delimiter появляется в выходных данных, ему предшествует текущий символ escapechar. Если escapechar не будет установлен, то writer поднимет Error , если столкнётся с какими-либо символами, которые требуют экранирования.

Дает указание reader не выполнять специальную обработку знаков цитаты.

Модуль csv определяет следующее исключение:

exception csv. Error ¶

Вызывается любой функцией при обнаружении ошибки.

Диалекты и параметры форматирования¶

Для упрощения задания формата входных и выходных записей, параметры форматирования группируются в диалекты. Диалект — это подкласс Dialect класса, имеющий набор специфических методов и единственный validate() метод. Создавая объекты reader или writer , программист может определить строку или подкласс класса Dialect как параметр диалекта. В дополнение, или вместо, параметра dialect, программист может также определить отдельные параметры форматирования, у которых есть те же имена как атрибуты, определенный ниже для класса Dialect .

Диалекты поддерживают следующие атрибуты:

Односимвольная строка, используемая для отделения полей. По умолчанию ‘,’ .

Управляет тем, как сущности quotechar, появляющиеся внутри поля, должны самостоятельно закавычиваться. Когда True , символ удваивается. Когда False , escapechar — используется как префикс к quotechar. По умолчанию он True .

При выводе, если doublequote False и не установлен escapechar, вызывается Error , если quotechar найден в поле.

Односимвольная строка используемая writer, чтобы экранировать delimiter, если quoting установлен в QUOTE_NONE и quotechar, если doublequote — False . При чтении escapechar удаляет какое-либо особое значение со следующего символа. По умолчанию используется значение None , которое отключает экранирование.

Используемая строка используемая для завершения строки, произведенная writer . По умолчанию используется значение ‘\r\n’ .

В reader жёстко закодированы опознавательные символы ‘\r’ или ‘\n’ как конец строки и игнорирует lineterminator. Это поведение может измениться в будущем.

Одиносимвольная строка используемая для закавычивания полей, содержащих специальные символы, такие как delimiter или quotechar, или которые содержат символы новой строки. По умолчанию используется значение ‘»‘ .

Контролирует, когда кавычки должны генерироваться writer и распознаваться reader. Он может принимать любые константы QUOTE_* (см. раздел Содержание модуля ) и по умолчанию имеет значение QUOTE_MINIMAL .

При True , пробелы непосредственно следующие за delimiter, игнорируются. Значение по умолчанию — False .

Когда True , вызывает исключение Error на плохом входном CSV. По умолчанию — False .

Объекты Reader¶

Объекты Reader ( DictReader сущности и объекты, возвращаемые функцией reader() ) содержат следующие публичные методы:

Возвращает следующую строку итерабельного объекта reader в виде списка (если объект возвращен из reader() ) или dict (если это экземпляр DictReader ), распарсиваемого в соответствии с текущим диалектом. Обычно вы должны вызывать его как next(reader) .

Объекты Reader содержат следующие публичные атрибуты:

Описание диалекта только для чтения, использующего парсер.

Количество сток читаемых из итерируемого источника. Это не то же самое, что и количество возвращаемых записей, поскольку записи могут охватить несколько строк.

У объектов DictReader есть следующий публичный атрибут:

Если не передан в качестве параметра, создаваемому объекту, то этот атрибут инициализируется при первом доступе или при чтении первой записи из файла.

Writer объекты¶

У объектов Writer ( DictWriter сущности и объекты, возвращённые функцией writer() ), есть следующий публичные методы. row должен быть итератором строк или чисел для объектов Writer и словаря, отображающего имена полей в строки или числа (передав им сначала str() ) для объектов DictWriter . Обратите внимание, что комплексные числа записываются в окружении родителей. Это может вызвать некоторые проблемы для других программ, которые читают файлы CSV (при условии, что они вообще поддерживают комплексные числа).

csvwriter. writerow ( row ) ¶

Записать параметр row в файл объекта writer, отформатированному согласно текущему диалекту. Возвратит возвращаемое значение вызываемое write методом основного объекта файла.

Изменено в версии 3.5: Добавлена поддержка произвольных итераторов.

Написать все элементы в rows (итерируемый из объектов row, как описано выше) к объекту файла writer’a, отформатированному согласно текущему диалекту.

Объекты Writer содержат следующий публичный атрибут:

Описание диалекта, используемого Writer’ом только для чтения.

У объектов DictWriter есть следующий публичный метод:

Записать строку с именами полей (как определено в конструкторе) в объект файла писателя, отформатированному согласно текущего диалекта. Возвратит возвращаемое значение csvwriter.writerow() вызываемое и используемое внутри.

Добавлено в версии 3.2.

Изменено в версии 3.8: writeheader() теперь также возвращает значение, возвращённое csvwriter.writerow() методом, который он использует внутри.

Примеры¶

Самый простой пример чтения CSV файла:

Чтение файла с альтернативным форматом:

Соответствующий простейший пример записи:

Так как open() используется для открытия CSV файла для чтения, файл будет по умолчанию декодирован в юникоде использую системную кодировку по умолчанию (см. locale.getpreferredencoding() ). Чтобы декодировать файл, используя другую кодировку, используйте аргумент encoding для open:

То же самое относится к записи в отличной от системной кодировки по умолчанию: укажите аргумент encoding при открытии выходного файла.

Регистрация нового диалекта:

Немного более продвинутое использование читателя — ловящий и сообщающий об ошибках:

И в то время как модуль непосредственно не поддерживает парсинга строк, то он может легко это сделать:

Добавить комментарий

Ваш адрес email не будет опубликован. Обязательные поля помечены *