Row 58988

Row ID: 58988 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 58988 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

An ml engineer is a data scientist that knows how to write production code and understands the limits of distributed systems. Understands that 500 mb is enough data for a single worker node, understands partition mapping, broadcasting, scheduling; understands rdd or partitioned data frames, and can manage ml flow pipelines... And actually likes managing these systems. Imo a good ml engineer would probably also like refactoring crappy data science code for data scientists, with a knack for being able to translate pandas based algos into spark or dask distributed tasks. The whole occupation is only useful if you can track hyper parameters and how your data generation process is changing.

FieldValue
text An ml engineer is a data scientist that knows how to write production code and understands the limits of distributed systems. Understands that 500 mb is enough data for a single worker node, understands partition mapping, broadcasting, scheduling; understands rdd or partitioned data frames, and can manage ml flow pipelines... And actually likes managing these systems. Imo a good ml engineer would probably also like refactoring crappy data science code for data scientists, with a knack for being …
label r/datascience
dataType comment
communityName r/datascience
datetime 2024-05-23
username_encoded Z0FBQUFBQm5Lak1ZZHZRdFBPS094RjREaTIya0ZfeURtalJPVkdEdGZyY0ZLSloxUk05WjM0Z1phRi0tdUpPVExJRkx3T0VseVZ2TUlsSnRJMGlYMFFZTldWc3lsQ1lMRVE9PQ==
url_encoded Z0FBQUFBQm5Lak9uWlRkYi1EeE9oZ01sTGNtT2RXV2NvQnI5OFdxUjZFM2VQRi1ZRDlKTmYwZzZGSXNnWHhiSmdaWTNfeEJKUm03RFRUcW9ta09uNUZzUzFIRkhFUFVXc0NwOFo0S1cyOTFpVktrQkNVajlYLTFMdE1yMFRvUHJQWmN0NzdtWC1qbE9wbzJtN0lrVmhkdTcxMkpzSURpb0ZVeHJNTEowWHNBTzVqcmxoSVRHbkZ2el9SckJ6alFlOEdfNzBSRU10N1R2

Raw Record

{
  "text": "An ml engineer is a data scientist that knows how to write production code and understands the limits of distributed systems. Understands that 500 mb is enough data for a single worker node, understands partition mapping, broadcasting, scheduling; understands rdd or partitioned data frames, and can manage ml flow pipelines... And actually likes managing these systems. Imo a good ml engineer would probably also like refactoring crappy data science code for data scientists, with a knack for being able to translate pandas based algos into spark or dask distributed tasks. The whole occupation is only useful if you can track hyper parameters and how your data generation process is changing.",
  "label": "r/datascience",
  "dataType": "comment",
  "communityName": "r/datascience",
  "datetime": "2024-05-23",
  "username_encoded": "Z0FBQUFBQm5Lak1ZZHZRdFBPS094RjREaTIya0ZfeURtalJPVkdEdGZyY0ZLSloxUk05WjM0Z1phRi0tdUpPVExJRkx3T0VseVZ2TUlsSnRJMGlYMFFZTldWc3lsQ1lMRVE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9uWlRkYi1EeE9oZ01sTGNtT2RXV2NvQnI5OFdxUjZFM2VQRi1ZRDlKTmYwZzZGSXNnWHhiSmdaWTNfeEJKUm03RFRUcW9ta09uNUZzUzFIRkhFUFVXc0NwOFo0S1cyOTFpVktrQkNVajlYLTFMdE1yMFRvUHJQWmN0NzdtWC1qbE9wbzJtN0lrVmhkdTcxMkpzSURpb0ZVeHJNTEowWHNBTzVqcmxoSVRHbkZ2el9SckJ6alFlOEdfNzBSRU10N1R2"
}

Entry Information