Removing DUPLICATE rows in hive based on columns

You can do the following :

select col1,col2,dayid,marketid,max(createdate) as createdate
from tablename
group by col1,col2,dayid,marketid

This way you are grouping the data by all the columns except the data so if there are rows with the same values in these columns they will be in the same group, and then, just "choose" the createdate you want by using an aggregate function like max/min etc.

Removing DUPLICATE rows in hive based on columns

Tags:

Hive

Related

Recent Posts