Skip to content

Introduces a new LogicalType: FILE#585

Open
brkyvz wants to merge 12 commits into
apache:masterfrom
brkyvz:fileType
Open

Introduces a new LogicalType: FILE#585
brkyvz wants to merge 12 commits into
apache:masterfrom
brkyvz:fileType

Conversation

@brkyvz

@brkyvz brkyvz commented Jun 9, 2026

Copy link
Copy Markdown

Rationale for this change

Introduces a new type called File as a typed FileReference. The design document is here.

The motivation is as follows:

Unstructured data ingestion is getting extremely popular with the advances in Generative AI. 
Today, our only means of dealing with unstructured data is to store it as a binary blob inside Parquet, 
or point to files that exist in some object store with a string. These solutions fail to address these use 
cases, because of scalability, usability, and governance issues.

We would like to introduce a new logical type annotation in Parquet called “File” for storing a struct that 
contains a path reference to a file with additional metadata. This reference may be to a file that exists 
(or expected to exist) in storage at a given path. We’d like to define the minimum required list of fields 
that would allow a client to correctly read the referenced data. Any additional metadata can be optionally 
stored by engines and table formats as necessary adjacent to this type. 

What changes are included in this PR?

Introduces the specification for FileType.

Do these changes have PoC implementations?

Yes:

Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated

@emkornfield emkornfield left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks reasonable to me a few minor comments for clarification.

Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated

@etseidl etseidl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a few questions I have after reviewing the Rust implementation.

Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread LogicalTypes.md Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread LogicalTypes.md Outdated
Comment thread src/main/thrift/parquet.thrift
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment on lines +732 to +733
The referenced bytes are compressed with the same `CompressionCodec` as the one
specified for the `inline` column.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's make this clearer.

Suggested change
The referenced bytes are compressed with the same `CompressionCodec` as the one
specified for the `inline` column.
The bytes referenced by a self-reference are compressed with the same `CompressionCodec` as the one specified for the `inline` column.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we point out CompressionCodec can differ per page?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@alkis thoughts? Is there a reason that the data needs to be compressed additionally? Most formats may be optimal anyway.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason is because we are making the assumption that the type of data stored in the file column is homogeneous. I argue that's a good assumption. If this column contains text, there is little reason to assume it will be a string when small and an image when large. Ergo if the writer chose to compress the column it should compress the self-referenced packed blobs too.

I wanted to make all packed references compressed as well but there was some pushback for that with the argument that external may be referenced by different parquet files and it may choose its own compression scheme (or none at all). I can see arguments both ways - it is a tradeoff.

Making all packed references inherit the CompressionCodec makes the spec more consistent.
Making only self-references inherit the CompressionCodec is friendlier to shared external references.

We pick one of the two, it is a tradeoff.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we point out CompressionCodec can differ per page?

Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread src/main/thrift/parquet.thrift Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated

@etseidl etseidl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @brkyvz, this looks good to me now.

@sfc-gh-sgrafberger sfc-gh-sgrafberger left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @brkyvz! Looks good to me.

Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
dejankrak-db added a commit to dejankrak-db/delta that referenced this pull request Jul 9, 2026
…a-io#585)

Align with the latest apache/parquet-format#585:
- size must be set whenever offset is set; a self-reference (no path) must set
  offset, and therefore size. Drop the now-invalid [offset, EOF) and [0, size)
  self-reference modes and the external [offset, EOF) mode from the resolution
  table; add explicit invalid rows.
- Define "set" (present, non-null, non-empty for strings) and allow sparse
  group definitions (a group need only define the fields it uses); add an
  inline-only example group.
- Fields matched case-sensitively by name; field IDs "if they exist".
- Readers should ignore unknown checksum algorithms.

Follows the consistent prose intent of PR delta-io#585; note its resolution table still
lists an [offset, EOF) row that contradicts its own validation section (offset
requires size) -- to be raised on the Parquet PR.

Co-authored-by: Isaac
@brkyvz

brkyvz commented Jul 10, 2026

Copy link
Copy Markdown
Author

Addressed your comments, @rok @pitrou please take a look. There's only one unaddressed comment on the CompressionCodec, which I will address after talking to @alkis

Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md
* A self-reference (`path` not set) must set `offset`. A value with neither `path` nor
`offset` set (and not `inline`) does not resolve and is invalid.
* `size` must be set whenever `offset` is set. A value that sets `offset` without `size`
is invalid. Because a self-reference must set `offset`, it must also set `size`.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How should implementations behave if offset is set but size is not? Are they required to fail?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally fail at write time, but is there a precedent for enforcing that? From your past comments it seems like we don't necessarily enforce contents of columns. Invalid files can be treated as null files in my opinion if we can't enforce it

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A "must" requirement is sufficient for writers. I think that requirement implies a failure at write time because otherwise the requirements are violated.

I'm more concerned about read time, which is why I'm asking. If this is invalid, what is the expected reader behavior? In the offset=null case, size is understood to be the length of the file. That doesn't seem right, but my guess is that readers will produce bytes up to the end of the file unless you specifically state that this should fail.

Comment thread LogicalTypes.md

@rdblue rdblue left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I left a few comments for clarification, but I think that this is ready either way.

Comment thread LogicalTypes.md
Comment thread LogicalTypes.md
Comment thread LogicalTypes.md
Comment thread LogicalTypes.md Outdated
the current file, so a file containing self-references is renamed or relocated as a
single unit.

The bytes referenced by a self-reference are compressed with the same `CompressionCodec`

@wgtmac wgtmac Jul 14, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The does not look right. Same column can have different codecs across different row groups. V2 pages can even decide whether to compress individually. Even if all row groups of the same column use the same codec, how can we tell whether self-referenced data is compressed or not?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reworded to

The bytes referenced by a self-reference use the `CompressionCodec`
defined by the `inline` column chunk's `ColumnMetadata`.

is that better @wgtmac ?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may have missed the relevant discussion on this so please correct me if I was wrong.

CompressionCodec defined by a ColumnMetadata is to compress all data pages of that column chunk in that specific row group. The description here means that we enforce all inline blocks referenced by this FILE column to always use the same CompressionCodec in the same row group? Is this too restrictive? Why can't we choose them to be uncompressed (like bloom filter) or compressed by a different scheme?

For other non-self-referenced cases, we can deduce compression codec based on content_type field, right? Shouldn't we do the same thing?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For other non-self-referenced cases, we can deduce compression codec based on content_type field, right? Shouldn't we do the same thing?

The intent of the field is to capture the content type of the data for filters like "find all PNG images". I don't expect this to describe compression because that wouldn't be very useful (all application/zip, for example).

I think the content_type field would behave like the HTTP Content-Type header that is the original media type before compression. HTTP uses a separate Content-Encoding header for compression.

This leaves the question of how to determine the equivalent of Content-Encoding for external references. We should probably spend some time considering that more.

The description here means that we enforce all inline blocks referenced by this FILE column to always use the same CompressionCodec in the same row group? Is this too restrictive? Why can't we choose them to be uncompressed (like bloom filter) or compressed by a different scheme?

Yes, I think that reading is accurate. Looks like this is a way to determine compression, but only for blobs that are located within the Parquet file itself as a self-reference.

I don't think that this is too restrictive since generic compression is typically uniform for files. I don't expect people to customize the compression used for these. We want compression to be possible (mandating uncompressed doesn't seem like a good choice), but we don't need to track it per value.

Comment thread LogicalTypes.md Outdated
```

Because every field is optional, a group need only define the fields it uses. A group
whose values are always stored inline may define just `inline`:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why defined content_type below but here says just inline? Seems inconsistent.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't see this is addressed. At least inline_file struct below still defines content_type.

Comment thread LogicalTypes.md
Comment thread LogicalTypes.md Outdated
Comment thread LogicalTypes.md Outdated
@brkyvz

brkyvz commented Jul 17, 2026

Copy link
Copy Markdown
Author

@rdblue @wgtmac @mapleFU @RussellSpitzer Thank you for the feedback! Addressed your comments. Hopefully things are clearer now! Let me know if they aren't

Comment thread LogicalTypes.md
| Algorithm | Encoding | Notes |
|-----------|---------------|----------------------------------------------------------|
| `ETAG` | opaque | the object-store eTag, not recomputable |
| `MD5` | lowercase hex | as defined in RFC 6151 represented as 32 hex characters |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would propose we reference RFC 1321 (The MD5 Message-Digest Algorithm) over RFC 6151 (Updated Security Considerations for the MD5 Message-Digest and the HMAC-MD5 Algorithms).

Suggested change
| `MD5` | lowercase hex | as defined in RFC 6151 represented as 32 hex characters |
| `MD5` | lowercase hex | as defined in RFC 1321 represented as 32 hex characters |

Comment thread LogicalTypes.md
Comment on lines +713 to +714
| `CRC32` | lowercase hex | as defined in RFC 3385, represented as 8 hex characters |
| `CRC32C` | lowercase hex | as defined in RFC 9260, represented as 8 hex characters |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These might fit better - I am in no way an expert on CRC32(c) RFC 2083, RFC 3385, but we would want to specify exact alghoritms used.

Suggested change
| `CRC32` | lowercase hex | as defined in RFC 3385, represented as 8 hex characters |
| `CRC32C` | lowercase hex | as defined in RFC 9260, represented as 8 hex characters |
| `CRC32` | lowercase hex | as defined in RFC 2083, represented as 8 hex characters |
| `CRC32C` | lowercase hex | as defined in RFC 3385, represented as 8 hex characters |

Comment thread LogicalTypes.md
`offset` set (and not `inline`) does not resolve and is invalid.
* `size` must be set whenever `offset` is set. A value that sets `offset` without `size`
is invalid. Because a self-reference must set `offset`, it must also set `size`.
* If `inline` is set, it supplies the bytes readers; producers may treat `inline` and the

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
* If `inline` is set, it supplies the bytes readers; producers may treat `inline` and the
* If `inline` is set, it supplies the bytes; producers may treat `inline` and the

Comment thread LogicalTypes.md
Comment on lines +740 to +748
| set | – | – | – | the inline bytes |
| – | set | – | – | whole external file at `path` |
| – | set | set | - | invalid |
| – | set | – | set | external `path`, `[0, size)` |
| – | set | set | set | external `path`, `[offset, offset + size)` |
| – | - | set | - | invalid |
| – | - | - | set | invalid |
| – | – | set | set | this file, `[offset, offset + size)` (self-reference) |
| – | – | – | – | nothing — invalid |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: this table uses both dash (-) and em dash (–) to indicate null input. Let's just use dash?.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.