• Wow.

    I’ve been processing a couple of billion rows of data on my machine, the fans didn’t even come on. WTF are they teaching “experts” these days, or has Elmo only hired people who claim that they can “wrangle data” and say “yes” ?

    • bleistift2 ( bleistift2@sopuli.xyz ) 
      link
      fedilink
      English
      arrow-up
      135
      ·
      2 years ago

      Even if querying data was processing-heavy and even if somehow the ‘hard drive’ got warm during this, then there still would need to be a hardware defect in order for the drive to overheat.

    • has Elmo only hired people who claim that they can “wrangle data” and say “yes” ?

      There’s two issues going on:

      1. Elmo’s sociopathic approach to laying people off is public knowledge, and top experts have the luxury of not even applying for his jobs.
      2. Elmo’s ability to judge engineering talent has likely been wildly exaggerated thanks to how he has successfully bought organizations full of talented people, in the past.
    • 60k rows is generally very usable with even wide tables in row formats.

      I’ve had pandas work with 1M plus rows with 100 columns in memory just fine.

      After 1M rows move on to something better like Dask, polars, spark, or literally any DB.

      The first thing I’d do with whatever data they’re running into issues with is rewrite it as partitioned and sorted parquet.

  • rumba ( rumba@lemmy.zip ) 
    link
    fedilink
    English
    arrow-up
    75
    ·
    2 years ago

    Unless I’m misreading it which is possible it’s awfully late, he said he processed 60,000 rows didn’t find what he was looking for but his hard drive overheated on the full pass.

    Discs don’t overheat because there was load. Even if he f***** up and didn’t index the data correctly (I assume it’s a relational database since he’s talking about rows) The disc isn’t just going to overheat because the job is big. It’s going to be lack of air flow or lack of heatsink.

    I guarantee you he was running on an external NVMe, and one of those little shitty-ass Chinese enclosures. Or maybe one of those self immolating SanDisk enclosures. Hell, maybe he’s on a desktop and he slept a raw NVMe on his motherboard without a heatsink

    There are times when you want a brilliant college student on your team, But you need seasoned professionals to help them through the things they’ve never seen before and never done before.

  • LillyPip ( LillyPip@lemmy.ca ) 
    link
    fedilink
    arrow-up
    59
    ·
    2 years ago

    This cannot be real, wtf. This is cartoon levels of ineptitude.

    Or sabotage by someone heading out? Please let this be resistance sabotage they haven’t noticed yet.

  • darkpanda ( darkpanda@lemmy.ca ) 
    link
    fedilink
    arrow-up
    30
    ·
    2 years ago

    What is this, a table for ants? Because that’s the average number of ants in an ant colony and it’s nowhere near an impressive amount of rows to be doing any sort of processing on. It wouldn’t be an impressive amount of rows if your rig was an i386DX-33 running off a 5” floppy.

    • 1rre ( 1rre@discuss.tchncs.de ) 
      link
      fedilink
      arrow-up
      12
      ·
      2 years ago

      Exactly, 60k rows is negligible enough in most cases that you can just treat it as free unless you’re doing a cross join on it or something, unless he’s doing something like using an unordered text file as his database with no ram or cache

  • golden_zealot ( golden_zealot@lemmy.ml ) 
    link
    fedilink
    English
    arrow-up
    27
    ·
    2 years ago

    I used to perform data analysis of robotics firmware logs which would generate several million log lines per hour and that was my second job out of college.

    I don’t know how you fuck up 60k lines that bad. Is he nesting 150 for loops and loading a copy of the data set in each one while mining crypto??

        • ButtDrugs ( ButtDrugs@lemm.ee ) 
          link
          fedilink
          arrow-up
          5
          ·
          2 years ago

          Storing large volumes of a text in a database column without optimization, then searching for small strings within it. It causes the database to basically search character by character to find a match by reading everything from disk. If you use indexes the database can do a lot of really incredible optimization to make finding values mich faster, and honestly string searching is better suited to a non-relational DB engine (which is why search engines don’t use relational DBs).

          Cartesian explosion is where you join related data together in a way that causes your result set to be wayyyy bigger than you expect. For example if you try to search through blog posts, but then also decide to bring in comments to search, then bring in the authors of those comments and all their comments from other posts. Result sets start to grow exponentially in that way, so maybe if you only search a few thousand blog posts you might be searching through millions of records because you designed your queries poorly.

        • manicdave ( manicdave@feddit.uk ) 
          link
          fedilink
          arrow-up
          4
          ·
          2 years ago

          If there’s something you want to search by in a database, you should index it.

          Indexing will create an ordered data structure that will allow much faster queries. If you were looking for the username gazter in an unindexed column, it would have to check literally every username entry. In a table of 1000000 entries it would check 1000000 times.

          In an indexed column it might do something like ask to be pointed to every name beginning with “g”, then of those ask to be pointed to every name with the second letter “a” and so on. It would find out where in the database gazter is by checking only six times.

          Substring matching is much more computationally difficult as it has to pull out each potentially matching value and run it through a function that checks if gazter exists somewhere in that value. Basically if you find yourself doing it you need to come up with a better plan.

          Cartesian explosion would be when your query ends up doing a shit load of redundant work. Like if the query to load this thread were to look up all the posters here, get all their posts, get the threads from those posts and filter on the thread id.

          • gazter ( gazter@aussie.zone ) 
            link
            fedilink
            arrow-up
            1
            ·
            2 years ago

            That’s very clear, thanks.

            I’m guessing you’d have to search the database to make the index, right? To search for ‘gazter’ you’d have had to go over the whole dataset and assigned each entry with a starting letter value, and so on?

            • manicdave ( manicdave@feddit.uk ) 
              link
              fedilink
              arrow-up
              2
              ·
              2 years ago

              When it comes to searching the database, the index will have already been created. When you create an index, it might take a while as the database engine reads all the data and creates a structure to shadow it. Each engine is probably different and I don’t know if any work exactly like that, but it’s an intuitive way to understand the basics of how B-trees work. You don’t really need to think much about how it works, just that if you want to use a column as a filter, you want to index it.

              However, when you’re thinking about the structure of a database it’s a good idea to think what you’ll want to do with it before hand and how you’ll structure queries. Sometimes searching columns without an index is unavoidable and then you’ve got to come up with other tricks to speed up your search. Like your doctor might find you (i’m presuming gaz is sort for gary and/or gareth here) with a query like SELECT * FROM patients WHERE birthdate = "01-01-1980" AND firstname LIKE "gar%" The db engine will first filter by birthdate which will massively reduce the amount of times it has to do the more intensive LIKE operation.

  • adarza ( adarza@lemmy.ca ) 
    link
    fedilink
    English
    arrow-up
    24
    ·
    2 years ago

    just a lame-ass excuse for not finding whatever evidence they were looking for.

    elsewhere, some seeding was done.

    now they’ll do the ‘full’ data grab and ‘find’ what they were looking for.

  • 1984 ( 1984@lemmy.today ) 
    link
    fedilink
    arrow-up
    22
    ·
    2 years ago

    I bet a million bucks the harddrive didnt “overheat”.

    Its just someone who doesnt know anything about computer hardware.

    Its like me saying my car overheated if there is smoke coming out of it. I know nothing about cars.