NOTE
Before creating request for Communishift namespace read our Communishift User Guide
Project Name: hatlas-elt
Project Administrators: hankuoffroad
Requires Persistent Storage?:yes
How much space do you need?: 50GB
Objective: Take an alternative approach for data engineering to build ELT (Extract, Load, Transform) pipeline with OLake
OLake is an open-source data replication and ELT (Extract, Load, Transform) framework designed to move data efficiently from various operational databases into open data lakehouse formats like Apache Iceberg.
OLake architecture https://olake.io/blog/olake-architecture/
Working group: Data WG Unofficial project name: Hatlas Project lead: Michael Winters Source data: datanommer, Postgresql Target output format: Parquet Silver / Iceberg tables
How I can contribute
Reference: Bronze/Silver pipeline test using an alternative ingestion data stack like olake https://olake.io/docs/connectors/postgres/setup/local/
The OLake PostgreSQL Source connector is compatible with Postgres 13. https://olake.io/docs/connectors/postgres/ https://olake.io/docs/understanding/compatibility-engines/
Size of a current DB copy (datanommer2): 1.4T
Note: I plan to test the Olake stack in an independent environment, much like a software fork, ensuring zero impact on the Data Lake infrastructure/the original project Michael Winters envisages. My goal is to document all findings and procedural steps clearly.
By submitting this ticket you agree to you have read and understood https://docs.fedoraproject.org/en-US/infra/communishift/#_service_usage_requirements
To be clear: my name is on this but it is not my request.
Hank is welcome to pursue this as one facet of exploring our Data WG options, but my understanding of https://docs.fedoraproject.org/en-US/infra/communishift/#_service_usage_requirements (and from prior infra Matrix threads) is that we cannot put any such data into Communishift.
4 Do not store or handle personal data on Communishift instance
Setting this as low / low ... but tend to assume mwinters is correct and this will be rejected.
Metadata Update from @james: - Issue priority set to: Waiting on Assignee (was: Needs Review) - Issue tagged with: low-gain, low-trouble
It's not clear to me... this is just an investigation? Or you plan to import/copy the datanommer db into it? Does it need storage for all that? Does it need anything aside an openshift project?
This issue has been migrated to Fedora Forge: https://forge.fedoraproject.org/infra/tickets/issues/12982
Please continue any further discussion there.
Metadata Update from @ryanlerch: - Issue close_status updated to: Migrated to Fedora Forge - Issue status updated to: Closed (was: Open)