Skip to content

Feature/mdf data sources - #69

Open
jonathanb-db wants to merge 5 commits into
mainfrom
feature/mdfDataSources
Open

Feature/mdf data sources#69
jonathanb-db wants to merge 5 commits into
mainfrom
feature/mdfDataSources

Conversation

@jonathanb-db

Copy link
Copy Markdown
Collaborator

Summary

Adds 3 DataSources for MF4 binary file loading, required data readers as well as a quickstart guide ("QUICKSTART.md" in the data_sources subdirectory). This is the first of many planned data sources to be added to impulse. The introduced readers are optimized to work with spark and use a stripe-memory reading approach to optimize IO-calls and general throughput compared to other solutions available. Please review the "README.md" in the code directory for details.

Currently this feature is experimental, since some parts of the mf4 standard are not yet implemented, missing functionality shall be added in the future once required. For a full list of current limitations review "KNOWN_LIMITATIONS.md" in the sub directory.

A Databricks Industry Solution that uses these data sources for a end-to-end ingestion pipeline is coming soon.

Changes

  • Added 3 Data sources to load mf4 files into a spark data frame
  • Added required readers to interpret mf4

Test Plan

  • Unit tests added/updated
  • Manual testing completed
  • Documentation updated (if applicable)

Checklist

  • Code follows project style guidelines
  • Self-review completed
  • No new linter warnings introduced

@jonathanb-db
jonathanb-db requested a review from a team as a code owner July 30, 2026 06:28
Comment thread uv.lock
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/ee/67/531ea369ba64dcff5ec9c3402f9f51bf748cec26dde048a2f973a4eea7f5/annotated_types-0.7.0.tar.gz", hash = "sha256:aff07c09a53a08bc8cfccb9c85b05f1aa9a2a6f23728d790723543408344ce89", size = 16081, upload-time = "2024-05-20T21:33:25.928Z" }
source = { registry = "https://pypi-proxy.cloud.databricks.com/simple/" }
sdist = { url = "https://pypi-proxy.cloud.databricks.com/packages/ee/67/531ea369ba64dcff5ec9c3402f9f51bf748cec26dde048a2f973a4eea7f5/annotated_types-0.7.0.tar.gz", hash = "sha256:aff07c09a53a08bc8cfccb9c85b05f1aa9a2a6f23728d790723543408344ce89", upload-time = "2024-05-20T21:33:25.928Z" }

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The databricks pypi-proxy is not reachable from the GH runner and this might be the reason why the lint check fails. Please check the uv.lock file of Impulse in the main branch: https://github.com/databrickslabs/impulse/blob/main/uv.lock

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants