Wow.
I’ve been processing a couple of billion rows of data on my machine, the fans didn’t even come on. WTF are they teaching “experts” these days, or has Elmo only hired people who claim that they can “wrangle data” and say “yes” ?
Even if querying data was processing-heavy and even if somehow the ‘hard drive’ got warm during this, then there still would need to be a hardware defect in order for the drive to overheat.
He hired a bunch of 19-25 year old. Not experts
Hey! Thats offensive to 19-25 year olds, there are many who just finished college/university and are more than aware.
They’re just role playing like in movies, with no idea of the consequences.
How on earth is it offensive to say they’re “not experts”? They’re not prodigies with PhDs. These specific young men are just technical enough and ideologically aligned.
Except they’re not, as you will know their tweet would be false after your first year of any technical (IT oriented) education.
First year? That shit is like A+ cert level knowledge or below, and A+ is damn near worthless. They would know that in the first few hours of a study guide
I was being generous when you consider the people in school who somehow pass, even when they don’t know a thing 🥲
Technical enough to be hired, is all I meant. 🙄
Apologies, if I came over as hostile. I did not get your meaning through text.
lol a 19-21 yo isnt going to have a degree lol,
There is nothing wrong with being 19-25. There’s something wrong with being wholly incompetent.
There’s not really anything wrong with being incompetent, so long as you have the humility to admit it and learn from people who know better, and try not to cause harm. That’s not Musk’s minions though.
I think it’s important to differentiate incompetence from ignorance. Ignorance is not knowing. Incompetence is not being able to fulfill the requirements for your assigned task. If you cannot fulfill the requirements for your given task, then you should not be given said task.
has Elmo only hired people who claim that they can “wrangle data” and say “yes” ?
There’s two issues going on:
- Elmo’s sociopathic approach to laying people off is public knowledge, and top experts have the luxury of not even applying for his jobs.
- Elmo’s ability to judge engineering talent has likely been wildly exaggerated thanks to how he has successfully bought organizations full of talented people, in the past.
60k rows is generally very usable with even wide tables in row formats.
I’ve had pandas work with 1M plus rows with 100 columns in memory just fine.
After 1M rows move on to something better like Dask, polars, spark, or literally any DB.
The first thing I’d do with whatever data they’re running into issues with is rewrite it as partitioned and sorted parquet.
My go-to tool of late is
duckdb, comes with binaries for most platforms, works out of the box, loads any number of database formats and is FAST.
60k isn’t that much, I frequently run scripts against multiple hundreds of thousands at work. Wtf is he doing? Did he duplicate the government database onto his 2015 MacBook Air?
deleted by creator
Sqlite can easily handle millions of rows. Don’t sell it short
How about a 6.4TB sqlite database?
Should be enough to hold 60k rows
I have an sqlite db that is a few GB in size, game saves using the format. Sadly almost all blob data, would love to play with it if it was a bit more readable
deleted by creator
60k is single json file
A TI-86 can query 60k rows without breaking a sweat.
If his hard drive overheated from that, he is doing something very wrong, very unhygienic, or both.
He probably mining crypto on top of running his SQL queries.
What? You don’t run your hard drives in the oven while baking brownies? It makes them zesty.
There must be more join statements than column names
Don’t know what Elmos minions are doing, but I’ve written code at least equally unefficient. It was quite a few years ago (the code was in written in perl) and I at least want to think that I’m better now (but I’m not paid to code anymore). The task was to pull in data from a CSV (or something like that, as I mentioned, it’s been a while) and it needed conversion to XML (or something similar).
The idea behind my code was that you could just configure which fields you want from arbitary source data and on where to place them on the whatever supported destination format. I still think that the basic idea behind that project is pretty neat, just throw in whatever you happen to have and have something completely else out of the other end. And it worked as it should. It was just stupidly hungry for memory. 20k entries would eat up several gigabytes of memory from a workstation (and back then it was premium to have even 16G around) and it was also freaking slow to run (like 0.2 - 0.5 seconds per entry).
But even then I didn’t need to tweet that my hard drive is overheating. I well understood that my code is just bad and I even improved it a bit here and there, but it was still so very slow and used ridiculous amounts of RAM. The project was pretty neat and when you had few hundred items to process at a time it was even pretty good, there was companies who relied on that code and paid for support. It just totally broke down with even a slightly bigger datasets.
But, as I already mentioned, my hard drive didn’t overheat on that load.
I’ve run searches over 60k lines of raw JSON on a 2015 MacBook air without any problems.
No, its an external drive, appearently.
I mean if we were to sort of steelman this thing, there sure can be database relations and queries that hit only 60k rows but are still hteavy as fuck.
My fucking events table of my synapse DB in postgres is nearly ten times as large, and I ported that from sqlite no long ago, in a matter of minutes. All of the data is on a 2*3 cluster of old 256GB SSDs, equaling about 1.5TB with Raid 0. That’s neither really fast, nor cool. But stable.
deleted by creator
Unless I’m misreading it which is possible it’s awfully late, he said he processed 60,000 rows didn’t find what he was looking for but his hard drive overheated on the full pass.
Discs don’t overheat because there was load. Even if he f***** up and didn’t index the data correctly (I assume it’s a relational database since he’s talking about rows) The disc isn’t just going to overheat because the job is big. It’s going to be lack of air flow or lack of heatsink.
I guarantee you he was running on an external NVMe, and one of those little shitty-ass Chinese enclosures. Or maybe one of those self immolating SanDisk enclosures. Hell, maybe he’s on a desktop and he slept a raw NVMe on his motherboard without a heatsink
There are times when you want a brilliant college student on your team, But you need seasoned professionals to help them through the things they’ve never seen before and never done before.
Can’t be a relational database, Musk said the government doesn’t use SQL.
Lol he also said cybertrucks don’t suck ;)
deleted by creator
music theaters also have rows, and they run on sql so logic checks out.
This cannot be real, wtf. This is cartoon levels of ineptitude.
Or sabotage by someone heading out? Please let this be resistance sabotage they haven’t noticed yet.
You guys arent running your software off raspberry pi’s with sdcards from the gas station?
My allowance is 5$ a month!
Look, all I’m saying is give Pis a chance.
deleted by creator
heh

“YOU’RE JUST JEALOUS” is such a fucking pussy-ass response, too.
Molly White is very bright, and she makes them feel inadequate so they “have to” attack her. It’s truly pathetic.
But her last name is White so it’s a real dilemma for them.
When the only thing that is stopping kids from dismantling your government is an O(N^N) algorithm
Are you telling me there’s a difference between an inner and a cross join?
Cross join is obviously faster, I don’t even have to write “on”
60k lol.
I regularly work with data in the 16tb range and weirdly my computer is fine. Git gud, doge scrubs.
Maybe
Maybe
They are just making shit up and doing jack shit
What is this, a table for ants? Because that’s the average number of ants in an ant colony and it’s nowhere near an impressive amount of rows to be doing any sort of processing on. It wouldn’t be an impressive amount of rows if your rig was an i386DX-33 running off a 5” floppy.
Exactly, 60k rows is negligible enough in most cases that you can just treat it as free unless you’re doing a cross join on it or something, unless he’s doing something like using an unordered text file as his database with no ram or cache
Buddy’s probably running code he got from GitHub Copilot that is used to do a visualization of a bubble sort for learning purposes.
I used to perform data analysis of robotics firmware logs which would generate several million log lines per hour and that was my second job out of college.
I don’t know how you fuck up 60k lines that bad. Is he nesting 150 for loops and loading a copy of the data set in each one while mining crypto??
Substring searches in unindexed large string columns or cartesian explosion caused by shitty joins would be my initial guess.
Largely ignorant, but data-curious person here.
…what?
Storing large volumes of a text in a database column without optimization, then searching for small strings within it. It causes the database to basically search character by character to find a match by reading everything from disk. If you use indexes the database can do a lot of really incredible optimization to make finding values mich faster, and honestly string searching is better suited to a non-relational DB engine (which is why search engines don’t use relational DBs).
Cartesian explosion is where you join related data together in a way that causes your result set to be wayyyy bigger than you expect. For example if you try to search through blog posts, but then also decide to bring in comments to search, then bring in the authors of those comments and all their comments from other posts. Result sets start to grow exponentially in that way, so maybe if you only search a few thousand blog posts you might be searching through millions of records because you designed your queries poorly.
If there’s something you want to search by in a database, you should index it.
Indexing will create an ordered data structure that will allow much faster queries. If you were looking for the username gazter in an unindexed column, it would have to check literally every username entry. In a table of 1000000 entries it would check 1000000 times.
In an indexed column it might do something like ask to be pointed to every name beginning with “g”, then of those ask to be pointed to every name with the second letter “a” and so on. It would find out where in the database gazter is by checking only six times.
Substring matching is much more computationally difficult as it has to pull out each potentially matching value and run it through a function that checks if gazter exists somewhere in that value. Basically if you find yourself doing it you need to come up with a better plan.
Cartesian explosion would be when your query ends up doing a shit load of redundant work. Like if the query to load this thread were to look up all the posters here, get all their posts, get the threads from those posts and filter on the thread id.
That’s very clear, thanks.
I’m guessing you’d have to search the database to make the index, right? To search for ‘gazter’ you’d have had to go over the whole dataset and assigned each entry with a starting letter value, and so on?
When it comes to searching the database, the index will have already been created. When you create an index, it might take a while as the database engine reads all the data and creates a structure to shadow it. Each engine is probably different and I don’t know if any work exactly like that, but it’s an intuitive way to understand the basics of how B-trees work. You don’t really need to think much about how it works, just that if you want to use a column as a filter, you want to index it.
However, when you’re thinking about the structure of a database it’s a good idea to think what you’ll want to do with it before hand and how you’ll structure queries. Sometimes searching columns without an index is unavoidable and then you’ve got to come up with other tricks to speed up your search. Like your doctor might find you (i’m presuming gaz is sort for gary and/or gareth here) with a query like
SELECT * FROM patients WHERE birthdate = "01-01-1980" AND firstname LIKE "gar%"The db engine will first filter by birthdate which will massively reduce the amount of times it has to do the more intensive LIKE operation.
60k rows of anything will be pulled into the file cache and do very little work on the drive. Possibly none after the first read.
just a lame-ass excuse for not finding whatever evidence they were looking for.
elsewhere, some seeding was done.
now they’ll do the ‘full’ data grab and ‘find’ what they were looking for.
Yeah. Hiring inexperienced children into government isn’t fraud, by itself. But I bet it makes fraud way easier.
I bet a million bucks the harddrive didnt “overheat”.
Its just someone who doesnt know anything about computer hardware.
Its like me saying my car overheated if there is smoke coming out of it. I know nothing about cars.
They even doubled down: https://xcancel.com/DataRepublican/status/1900565922831618202#m

Well thanks for trying to keep us from catching it I guess
thank god, i am going to be ok, i was worried it might be contagious (it’s not, he just has brain damage)
Can you screenshot or something? I can’t load that link
Does this one work? https://nitter.net/DataRepublican/status/1900565922831618202#m Its a bit too long to screenshot.
That works, thanks
So it was a probably shitty external HD… that’s a whole other can of worms


















Good lord.

