August 22, 2026

Preparando databricks certified engineer professional (V)

Seguimos repasando y estudiando, ha habido drama con Snowflake, así que tengo más motivación. Section 3: Data Transformation, Cleansing, and Quality Write efficient Spark SQL and PySpark code to apply advanced data transformations, including window functions, joins, and aggregations, to manipulate and analyze large Datasets. Aquí voy a hablar de las reglas de oro: a) Lee lo mínimo indispensable. Si hay particiones, úsalas, si no necesitas una columna, no la leas. b) Rompe el plan una vez. Hubo un tiempo que vi abusar del .cache(). La realidad es que partir las ejecuciones rara vez era mejor. Mucho más fiable confiar en el disk caché. c) Esto es programación funcional, componer es ganar. d) Ordena solo si lo necesitas, que no es gratis, leñe, que tiras un shuffle. Read more

August 21, 2026

Preparando databricks certified engineer professional (IV)

Esto es una pequeña maratón, cada día, un tema :) Hoy hablaré de dos cosas interesantes, de data modeling y cost & optimization. Cambio un poco el formato porque esto más bien, será un pequeño dump de mis memorias. Section 6: Cost & Performance Optimization Understand delta optimization techniques, such as deletion vectors and liquid clustering. Delta fue una gran ventaja sobre parquet, antes eliminar datos de una tabla, había que o reescribir la tabla entera, o al menos la partición de esta. Cuando llegó delta, esto se hacía automáticamente, muchísimo más cómodo, ¿pero te imaginas reescribir un fichero de mil millones de filas solo porque quieres eliminar 4 datos? Pues esto pasaba, al menos hasta que llegaron los deletion vectors Read more

August 18, 2026

Preparando databricks certified engineer professional (III)

Esto es una pequeña maratón, cada día, un tema :) Hoy hablaré de dos cosas interesantes, de data modeling y cost & optimization. Cambio un poco el formato porque esto más bien, será un pequeño dump de mis memorias. Section 6: Cost & Performance Optimization Understand delta optimization techniques, such as deletion vectors and liquid clustering. Delta fue una gran ventaja sobre parquet, antes eliminar datos de una tabla, había que o reescribir la tabla entera, o al menos la partición de esta. Cuando llegó delta, esto se hacía automáticamente, muchísimo más cómodo, ¿pero te imaginas reescribir un fichero de mil millones de filas solo porque quieres eliminar 4 datos? Pues esto pasaba, al menos hasta que llegaron los deletion vectors Read more

August 17, 2026

Preparando databricks certified engineer professional (II)

Continuamos aprendiendo, hoy con github caído :) Hoy me toca centrarme en Data Sharing and Federation y Data Governance. De nuevo vamos a la guía del PDF: Section 4: Data Sharing and Federation Demonstrate delta sharing securely between Databricks deployments using Databricks to Databricks Sharing (D2D) or to external platforms using the open sharing protocol (D2O). Configure Lakehouse Federation with proper governance across the supported source Systems. Use Delta Share to share live data from Lakehouse to any computing platform. <- Eso es ahora Opensharing, así que no sé. Ya pido perdón porque esta parte es un poco de peñazo teórico, pero es lo que hay. De paso dejo en negrita la típica pregunta trampa: Read more

August 16, 2026

Preparando databricks certified engineer professional (I)

Continuando en mi cambio de trabajo por año. Intento volver a databricks. Es la tecnología que mejor se adapta a mí y que me permite mantener casi sin esfuerzo esa mentalidad constante de kaizen. Hace unos meses fue el Databricks Learning Festival donde por completar unos cursos, obtuve un 50% de descuento en certificaciones. Ahora tengo el examen en aproximadamente 11 días y el problema, es que hace dos años que no toco databricks. Hace dos años mi stack se basaba en: Read more

January 26, 2024

Querying the databricks api

Exploring databricks SQL usage At my company, we adopted databricks SQL for most of our users. Some users have developed applications that use the JDBC connector, some users have built their dashboards, and some users write plain ad-hoc queries. We wanted to know what they queried, so we tried to use Unity Catalog’s insights, but it wasn’t enough for our case. We work with IOT and we are interested in what filters they apply within our tables. Read more

October 2, 2023

Repairing metadata unity catalog

I’ve been subscribed to https://www.dataengineeringweekly.com/p/data-engineering-weekly-148 for years. This last number included several on-call posts on Medium. I found these quite useful. Today, I got an alert from Metaplane that a cost monitor dashboard was out of date. I checked the processes, and everything was fine. I ran a query to check the freshness of the data and it was ok too. Metaplane checks our delta table freshness by querying the table information available in the Unity Catalog. For some unknown reason that metadata didn’t receive any update. I ran an optimization operation (the table is tiny) and the metadata didn’t update either. Read more

July 28, 2023

Adding extra params on DatabricksRunNowOperator

With the new Databricks jobs API 2.1 you have different parameters depending on the kind of tasks you have in your workflow. Like: jar_params, sql_params, python_params, notebook_params… And not always the airflow operator is ready to handle all of them. If we check the current release of the DatabricksRunNowOperator, we can see that there is only support for: notebook_params python_params python_named_parameters jar_params spark_submit_params And not the query_params mentioned earlier. But there is a way of combining both, there is a param called json that allows you to write the payload of a databricksrunnow and it will also merge the content of the JSON with your named_params! Read more

2017-2026 Adrián Abreu