civex.parse_table
Parse a CSV, TSV or other delimited file's bytes into a pandas DataFrame — the usual first step before civex.rows_to_records or civex.upsert_records.
Requires pandas, which ships with civex. It is imported inside
invoke(), not at module load, so this plugin registers quickly and validates without loading pandas.
Plugin ID: civex.parse_table
Category: data-sources
Capabilities: none
Config
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
delimiter |
string | no | , | Column separator: one character, or a name (comma, tab, semicolon, pipe, space). Used when delimiters isn't set. |
delimiters |
array | null | no | null | Several separators the file may use, such as [comma, tab] to read both CSV and TSV files. The one the file's own lines use is picked; a file using none of them is read with the first. Overrides delimiter. |
encoding |
string | no | utf-8 | Text encoding to decode the file with. |
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
bytes |
bytes | yes | Raw file contents: CSV, TSV or other delimited text. |
Outputs
| Name | Type | Required | Description |
|---|---|---|---|
table |
table | yes | The parsed rows and columns. |
- id: load
plugin: civex.load_file
config:
field: selection_table
- id: parse
plugin: civex.parse_table
inputs:
bytes: load.bytes
Reading CSV and TSV with one step
Give delimiters a list of the separators a file may use, and the step picks the
one the file actually uses, from its own lines (fields inside double quotes are
ignored). A name such as tab can stand for the character:
- The separator is chosen from the first lines: the one that splits every line into the same number of fields wins, then the one that appears most, then the one listed first. A file that uses none of them (a single column) is read with the first.
- Names:
comma,tab,semicolon,pipe,space,colon.\talso means tab. Each entry indelimitersmust be one character. delimiter(singular) still sets one separator, exactly as before, and accepts the same names. A multi-character value is still passed to pandas, which reads it as a regular expression.delimitersoverridesdelimiterwhen both are set.- A byte-order mark at the start of a UTF-8 file (common from Excel) is dropped, so it doesn't end up in the first column's name.