{"@attributes":{"version":"2.0"},"channel":{"title":"Bigquery on Or Elimelech","link":"https:\/\/or-e.net\/tags\/bigquery\/","description":"Recent content in Bigquery on Or Elimelech","image":{"title":"Or Elimelech","url":"https:\/\/or-e.net\/%3Clink%20or%20path%20of%20image%20for%20opengraph,%20twitter-cards%3E","link":"https:\/\/or-e.net\/%3Clink%20or%20path%20of%20image%20for%20opengraph,%20twitter-cards%3E"},"generator":"Hugo -- 0.151.0","language":"en","lastBuildDate":"Mon, 14 Mar 2022 00:00:00 +0000","item":{"title":"Build a datalake on top of BigQuery","link":"https:\/\/or-e.net\/posts\/bigquery-lakehouse\/","pubDate":"Mon, 14 Mar 2022 00:00:00 +0000","guid":"https:\/\/or-e.net\/posts\/bigquery-lakehouse\/","description":"<p>Google BigQuery is a very powerful, serverless data warehosue that\nlets you ingest unlimited data on a pay-per-use basis (storage + querying).\nThe primary advantage of data warehouses is the ability to quickly query and analyze immense amounts of <strong>structured<\/strong> data.<\/p>\n<p>Modern data warehouses support new, unstructured data types such as JSON, Avro, and so on, which makes these data warehouses a great contender for data lakes.\nBigQuery recently added native JSON column type, which we can leverage for our semi-structured lakehouse.\nWith JSON support we can skip the traditional datalakes that are just blob storage with files, from classic Hadoop + hive through GCS, S3, Athena etc.\nMaintaining such infrastructure and understanding the low-level components, such as metastore, ORC, and Parquet is an expensive process for most startups, financially and with regard to domain expertise.<\/p>"}}}