• I absolutely despise Firebase Firestore (the database technology that was “hacked”). It’s like a clarion call for amateur developers, especially low rate/skill contractors who clearly picked it not as part of a considered tech stack, but merely as the simplest and most lax hammer out there. Clearly even DynamoDB with an API gateway is too scary for some professionals. It almost always interfaces directly with clients/the internet without sufficient security rules preventing access to private information (or entire database deletion), and no real forethought as to ongoing maintenance and technical debt.

    A Firestore database facing the client directly on any serious project is a code smell in my opinion.

    • Ah yes, Firebase. The Google version of leaking all your company data through a public S3 bucket

      I remember when they launched and started pushing it in the Android dev community. Actually won a Google Pixel at a Firebase sponsored hackathon in my town…after that I never touched Firestore again. Using that ACL language to restrict access, you could see the massive foot gun from a mile away

  • JackbyDev ( JackbyDev@programming.dev ) 
    link
    fedilink
    English
    arrow-up
    23
    ·
    edit-2
    1 year ago

    Hack has at least two definitions in a computing context.

    1. A nifty trick or shortcut that is useful. “Check out this hack to increase your productivity.”
    2. Accessing something you shouldn’t. “They hacked into the database.”

    A lot of times they sort of get used in conjunction to describe interesting ways to gain access to secure systems, but using it to describe accessing insecure things you shouldn’t is still a valid usage of the phrase.

    That said I definitely wanna see the company face charges for this, this is insane.

      • Rivalarrival ( Rivalarrival@lemmy.today ) 
        link
        fedilink
        English
        arrow-up
        7
        ·
        1 year ago

        Terrible analogy. A webserver is not at all like a door. It doesn’t block or allow traffic to and from your file system.

        A web server is more like a receptionist. It handles requests. “Can I have your basic catalog?” “Certainly, here you go.”

        “Can I get this item from your basic catalog?” “Certainly.”

        “I don’t see it in your catalog, but my buddy said he got this other item from you. Can I have this other item too?” “Absolutely.”

        “Can I borrow your stapler?” Sure. “How about a pad of paper?” “Of Course”. “Can I just have the contents of your supply closet?” “Here you go.” “How about your accounting files, can I get those?” “No problem!” “How about your entire customer list?” “Consider it done!”

        When you hire a receptionist and specifically tell them to give customers anything they request, that’s entirely on you. You have to at least make a token effort to restrict access to only authorized users before you can even claim that a particular user was unauthorized.

        This wasn’t burglary. This was putting up signs that say “come in” and labeling everything in your house with “free” stickers.

      • @SpaceCowboy @JackbyDev

        In a legal context there’s also the concept of a “reasonable expectation of privacy”. The computer abuse and fraud act defines hacking as accessing data or systems you are not authorized to access.

        A better analogy is putting your journal in a public library and getting mad when somone reads it.

        I’m not saying what these ass holes did was right, I’m saying that the company weakened their legal position by not protecting the data.

        • Terrible analogy. You have permission to read books in a library.

          Forgetting to lock your door isn’t granting permission to people enter your house, and it doesn’t grant people permission to take your valuables. It may be neglectful to leave your door unlocked, but it doesn’t imply granting permission to enter your house.

          Same goes with computer security. Leaving your computer insecure may be neglectful, but it does not imply someone has permission to take your data.

          • @SpaceCowboy

            Then how do I know what I am not allowed to access?

            In this specific case there was no (formal) indication that the data was out of bounds.

            I can’t put 10 pdf files in a web dir and claim 5 are public and 5 are private, then charge you with a crime for viewing them.

            You can’t have “unauthorized access” when there’s no authorization at all

            • If I’m clicking around on a website and find a gallery of images, that’s something I’m supposed to have access to. If I start typing in URLs that aren’t linked anywhere on the site, then I’m accessing stuff the site hasn’t explicitly indicated I have access to. If I’m doing this with the intent of getting data and distributing to others, then yeah that would be illegal.

              The law allows for someone to exercise judgement. The people who do this are not so coincidentally called Judges. If the 4chan guys had have been white hat and reported the issue to the site owners, then they’d be fine. But it’s obvious to anyone their intent was to get private information, they poked around to find some private information, and then distributed that private information to others causing a privacy violation. Yes, it was easier to do than it should have been, but it’s obvious they had malicious intent and it’s obvious they were accessing information they weren’t supposed to access.

              A crime being really easy to commit doesn’t make it no longer a crime. Many times I’ve seen things that I could easily steal, but I don’t steal things when I have an opportunity to do so because a) stealing is wrong and b) saying “they just left this thing out there in a place anyone could steal it” would not be any kind of legal defense. Simply because you’re presented an opportunity to do a crime doesn’t mean it’s acceptable to do a crime, both legally and morally speaking.

              • Rivalarrival ( Rivalarrival@lemmy.today ) 
                link
                fedilink
                English
                arrow-up
                2
                ·
                1 year ago

                I start typing in URLs that aren’t linked anywhere on the site, then I’m accessing stuff the site hasn’t explicitly indicated I have access to.

                Doesn’t work like that. With the policy you describe, anyone who ever sees a “404” error is a criminal.

                I don’t have to publish everything I am willing to offer. You are free to ask for something I may or may not have. I get to decide how to respond to your request.

                To use your analogy, I can walk up to your door and request a glass of water. You’ve never explicitly offered a glass of water to anyone; I’m still allowed to ask. If you dont want me to have your water, you can say “No” or you can ignore me.

                When you go ahead and give me a glass of water, you don’t get to claim I stole it from you. It is not theft to ask.

                You have to make some sort of effort to have your web server limit my access, and I have to make some sort of effort to convince your webserver to bypass those restrictions before you can claim I am exceeding my authorization.

        • iii ( iii@mander.xyz ) 
          link
          fedilink
          English
          arrow-up
          6
          ·
          1 year ago

          A better analogy is putting your journal in a public library and getting mad when someone reads it.

          Good analogy indeed. I’d go one step further and add: it’s like promising others you’ll keep their diary safe, then putting it in a public library, to then get mad when someone reads it.

          • @iii

            Yeah the internet by design is a public space, and we must be responsible and treat it as such when handling sensative data.

            Again, it was very wrong for people to take that data and especially to post like that.

            The company also has to do their part and produce at least some kind of barrier to the data.

            Even using UUIDs and making sure the data wasn’t query-able would have been something.

      • JackbyDev ( JackbyDev@programming.dev ) 
        link
        fedilink
        English
        arrow-up
        7
        ·
        1 year ago

        Thank you! I feel like I’m taking crazy pills reading people’s reactions to this. And if it was a business instead of your house and it was customer data you weren’t protecting you should still be in trouble too. It’s like people think only one side can be in the wrong in this or that because the data wasn’t secured and in the public that gives them free reign to post it everywhere. I wonder how those people would feel if their addresses were leaked. Afterall, if you’re a homeowner your name is attached to the property and is publicly accessible.

      • JackbyDev ( JackbyDev@programming.dev ) 
        link
        fedilink
        English
        arrow-up
        5
        ·
        1 year ago

        It can be both. The company can be at fault for not keeping something secure while the people who steal the data are at fault for stealing data. Data leaks and hacks are not mutually exclusive.

        • percent ( percent@infosec.pub ) 
          link
          fedilink
          arrow-up
          1
          ·
          1 year ago

          I don’t disagree with your main point, but I’m not sure it’s really even “stealing”, as that means to take without permission. In this case, the storage permissions were configured so that the files were publicly available to everyone, so everyone had permission to access them.

          Semantics though. It’s still unethical to access that data, even if it’s not technically stealing.

  • jaybone ( jaybone@lemmy.zip ) 
    link
    fedilink
    English
    arrow-up
    14
    ·
    1 year ago

    What was the BASE_URL here? I’m guessing that’s like a profile page or something?

    So then you still first have to get a URL to each profile? Or is this like a feed URL?

  • Rhaedas ( Rhaedas@fedia.io ) 
    link
    fedilink
    arrow-up
    13
    ·
    1 year ago

    Even the best models fine tuned for coding still have training that was based on both good and bad examples of programming from humans. And since it’s not AGI but using probability to generate the code, you’re going to get crap programming logic dependent on how often such things were used and suggested by humans to other humans. Googling for an answer on how to code something pulls up all sorts of answers from many sources, but reading through them, many are terrible. An LLM doesn’t know that, it just knows that humans liked some answers better than others, so GIGO.

  • I wonder if their data is poisoned by below average Dev. I mean if your test subjects are met or below Dev and mad Ethel lost 20% efficiency imagine what you can do to good dev

    • Thorry84 ( Thorry84@feddit.nl ) 
      link
      fedilink
      arrow-up
      9
      ·
      edit-2
      1 year ago

      Not below average dev necessarily, but when posting code examples on the internet people often try to get a point across. Like how do I solve X? Here is code that solves X perfectly, the rest of the code is total crap, ignore that and focus on the X part. Because it’s just an example, it doesn’t really matter. But when it’s used to train an LLM it’s all just code. It doesn’t know which parts are important and which aren’t.

      And this becomes worse when small little bits of code are included in things like tutorials. That means it’s copy pasted all over the place, on forums, social media, stackoverflow etc. So it’s weighted way more heavily. And the part where the tutorial said: “Warning, this code is really bad and insecure, it’s just an example to show this one thing” gets lost in the shuffle.

      Same thing when an often used pattern when using a framework gets replaced by new code where the framework does a little bit more so the same pattern isn’t needed anymore. The LLM will just continue with the old pattern, even though there’s often a good reason it got replaced (for example security issues). And if the new and old version aren’t compatible with each other, you are in for a world of hurt trying to use an LLM.

      And now with AI slop flooding all of these places where they used to get their data, it just becomes worse and worse.

      These are just some of the issues why using an LLM for coding is probably a really bad idea.

      • floofloof ( floofloof@lemmy.ca ) 
        link
        fedilink
        arrow-up
        2
        ·
        1 year ago

        Yeah, once you get the LLM’s response you still have to go to the documentation to check whether it’s telling the truth and the APIs it recommends are current. You’re no better off than if you did an internet search and tried to figure out who’s giving good advice, or just fumbled your own way through the docs in the first place.

        • You’re no better off than if you did an internet search and tried to figure out who’s giving good advice, or just fumbled your own way through the docs in the first place.

          These have their own problems ime. Often the documentation (if it exists) won’t tell you how to do something, or it’s really buried, or inaccurate. Sometimes the person posting StackOverflow answers didn’t actually try running their code, and it doesn’t run without errors. There are a lot of situations where a LLM will somehow give you better answers than these options. It’s inconsistent, and the reverse is true also, but the most efficient way to do it is to use all of these options situationally and as backups to each other.

          • floofloof ( floofloof@lemmy.ca ) 
            link
            fedilink
            arrow-up
            2
            ·
            1 year ago

            Yes, it can be useful in leading you to look in the right place for more information, or orienting you with the basics when you’re working with a technology that’s new to you. But I think it wastes my time as often as not.

            • That’s implying that the quality of information from other sources is always better, but I’m saying that’s sometimes not true; when you’re trying to figure out the syntax for something, documentation and search engines have failed you, and the traditional next step would be to start contacting people or trying to find the answer in unfamiliar source code, sometimes a LLM can somehow just tell you the answer at that point and save the trouble. Of course you have to test that answer because more often than not it will just make up a fake one but that just takes a few seconds.

              There are some situations I’m going back to search engines as a first option though, like error messages, LLMs seem to like to get tunnel vision on the literal topic of the error, while search results will show you an unintuitive solution to the same problem if it’s a very common one.

        • Kay Ohtie ( kayohtie@pawb.social ) 
          link
          fedilink
          English
          arrow-up
          2
          ·
          1 year ago

          whether it’s telling the truth

          “whether the output is correct or a mishmash”

          “Truth” implies understanding that these don’t have, and because of the underlying method the models use to generate plausible-looking responses based on training data, there is no “truth” or “lying” because they don’t actually “know” any of it.

          I know this comes off probably as super pedantic, and it definitely is at least a little pedantic, but the anthropomorphism shown towards these things is half the reason they’re trusted.

          That and how much ChatGPT flatters people.

          • floofloof ( floofloof@lemmy.ca ) 
            link
            fedilink
            arrow-up
            2
            ·
            1 year ago

            Yeah, it has no notion of being truthful. But we do, so I was bringing in a human perspective there. We know what it says may be true or false, and it’s natural for us to call the former “telling the truth”, but as you say we need to be careful not to impute to the LLM any intention to tell the truth, any awareness of telling the truth, or any intention or awareness at all. All it’s doing is math that spits out words according to patterns in the training material.

            • Kay Ohtie ( kayohtie@pawb.social ) 
              link
              fedilink
              English
              arrow-up
              2
              ·
              1 year ago

              I figured and I know it’s shorthand, it’s my own frustration that said shorthand has partly enabled the anthropomorphism that it’s enjoyed.

              Leave the anthropomorphism to pets, plants, and furries, basically. And cars. It’s okay to call cars like that. They know what they did.